this post was submitted on 30 Sep 2023

1095 points (98.8% liked)

Open Source

36636 readers

213 users here now

All about open source! Feel free to ask questions, and share news, and interesting stuff!

Useful Links

Rules

Posts must be relevant to the open source ideology
No NSFW content
No hate speech, bigotry, etc

Related Communities

Community icon from opensource.org, but we are not affiliated with them.

founded 5 years ago

MODERATORS

[email protected]

1095

Mozilla.ai is a new startup and community funded with 30M from Mozilla that aims to build trustworthy and open-source AI ecosystem (mozilla.ai)

submitted 2 years ago by [email protected] to c/[email protected]

103 comments fedilink hide all child comments

top 50 comments

sorted by: hot top controversial new old

[–] [email protected] 157 points 2 years ago* (last edited 2 years ago) (3 children)

My mind immediately went to a horizon zero dawn like dystopia where the Mozilla AI is the only thing left protecting humans from various malevolent AIs bent on consuming the human race

[–] [email protected] 38 points 2 years ago (2 children)

Mozilla is Gaia, ChatGPT is hades?

[–] [email protected] 24 points 2 years ago* (last edited 2 years ago) (2 children)

I think by that point ChatGPT would be more like Apollo, keeping the knowledge of humanity. I feel like one of the more corporate AIs will go full HADES, I'm thinking Bard. It will get a mysterious signal from space that switches it's core protocol from "don't be evil" to "be evil."

[–] [email protected] 4 points 2 years ago

A “mysterious signal from space” being just the fact that it’s owned by Google

load more comments (1 replies)

[–] [email protected] 5 points 2 years ago

SPOOOOOILER

load more comments (2 replies)

[–] [email protected] 91 points 2 years ago

Incredibly welcomed. We need more ethical, non-profit AI researchers in the sea of corporate for-profit AI companies.

[–] [email protected] 53 points 2 years ago

I want to give them the benefit of the doubt. I really do. I am going to watch this with a critical eye, however.

[–] [email protected] 43 points 2 years ago (2 children)

I'll believe it when I see it.

I'm so goddamn tired of "open source" turning into subscription models restricting use cases because the company wants to appease conservative investors.

[–] [email protected] 37 points 2 years ago

Mozilla has a very strong track-record though. They've been around for a very long time, and have stuck to free open-source principles the whole time.

[–] [email protected] 4 points 2 years ago* (last edited 2 years ago) (1 children)

That's basically only OpenAI, maybe some obscure startups as well. Mozzila is far too old and niche to get away with that anyway.

load more comments (1 replies)

[–] [email protected] 36 points 2 years ago (2 children)

I would really like Mozilla to make the best browser in the world please.

[–] [email protected] 101 points 2 years ago (1 children)

they do

load more comments (1 replies)

[–] [email protected] 14 points 2 years ago

As a (very recently) former chrome user, they do already.

[–] [email protected] 31 points 2 years ago

Wishing they would say something more, the site has been like that for some time.

[–] [email protected] 26 points 2 years ago (8 children)

As much as I love Mozilla, I know they're going to censor it (sorry, the word is "alignment" now) the hell out of it to fit their perceived values. Luckily if it's open source then people will be able to train uncensored models

[–] [email protected] 72 points 2 years ago (29 children)

What in the world would an "uncensored" model even imply? And give me a break, private platforms choosing to not platform something/someone isn't "censorship", you don't have a right to another's platform. Mozilla has always been a principled organization and they have never pretended to be apathetic fence-sitters.

[–] [email protected] 40 points 2 years ago (5 children)

This is something I think a lot of people don't get about all the current ML hype. Even if you disregard all the other huge ethics issues surrounding sourcing training data, what does anybody think is going to happen if you take the modern web, a huge sea of extremist social media posts, SEO optimized scams and malware, and just general data toxic waste, and then train a model on it without rigorously pushing it away from being deranged? There's a reason all the current AI chatbots have had countless hours of human moderation adjustment to make them remotely acceptable to deploy publicly, and even then there are plenty of infamous examples of them running off the rails and saying deranged things.

Talking about an "uncensored" LLM basically just comes down to saying you'd like the unfiltered experience of a robot that will casually regurgitate all the worst parts of the internet at you, so unless you're actively trying to produce a model to do illegal or unethical things I don't quite see the point of contention or what "censorship" could actually mean in this context.

[–] [email protected] 16 points 2 years ago

It means they can’t make porn images of celebs or anime waifus, usually.

load more comments (4 replies)

[–] [email protected] 21 points 2 years ago

I fooled around with some uncensored LLaMA models, and to be honest if you try to hold a conversation with most of them they tend to get cranky after a while - especially when they hallucinate a lie and you point it out or question it.

I will never forget when one of the models tried to convince me that photosynthesis wasn't real, and started getting all snappy when I said I wasn't accepting that answer 😂

Most of the censorship "fine tuning" data that I've seen (for LoRA models anyway) appears to be mainly scientific data, instructional data, and conversation excerpts

[–] [email protected] 17 points 2 years ago (2 children)

There's a ton of stuff ChatGPT won't answer, which is supremely annoying.

I've tried making Dungeons and Dragons scenarios with it, and it will simply refuse to describe violence. Pretty much a full stop.

Open AI is also a complete prude about nudity, so Eilistraee (Drow godess that dances with a sword) just isn't an option for their image generation. Text generation will try to avoid nudity, but also stop short of directly addressing it.

Sarcasm is, for the most part, very difficult to do... If ChatGPT thinks what you're trying to write is mean-spirited, it just won't do it. However, delusional/magical thinking is actually acceptable. Try asking ChatGPT how licking stamps will give you better body positivity, and it's fine, and often unintentionally very funny.

There's plenty of topics that LLMs are overly sensitive about, and uncensored models largely correct that. I'm running Wizard 30B uncensored locally, and ChatGPT for everything else. I'd like to think I'm not a weirdo, I just like D&d... a lot, lol... and even with my use case I'm bumping my head on some of the censorship issues with LLMs.

load more comments (2 replies)

load more comments (26 replies)

[–] [email protected] 6 points 2 years ago

As an aside I'm in corporate. I love how gung ho we are on AI meanwhile there are lawsuits and potential lawsuits and investigative journalism coming out on all the shady shit AI and their companies are doing. Meanwhile you know the SMT ain't dumb they know about all this shit and we are still driving forward.

load more comments (6 replies)

[–] [email protected] 18 points 2 years ago* (last edited 2 years ago)

I remember a time when open-source software was developed without a pre-order business model.

"This new company will be led by Managing Director Moez Draief. Moez has spent over a decade working on the practical applications of cutting-edge AI as an academic at Imperial College and LSE, and as a chief scientist in industry. Harvard’s Karim Lakhani, Credo’s Navrina Singh and myself will serve as the initial Board of Mozilla.ai. "

This money will 90% go to paying the salary of these managers.

[–] [email protected] 18 points 2 years ago (2 children)

All I want to know is if they are going to pillage people's private data and steal their creative IP or not.

Ethical AI starts and ends with open, transparent, legitimate and ethically sourced training data sets.

[–] [email protected] 16 points 2 years ago (5 children)

Using copyrighted material for research is fair use. Any model produced by such research is not itself a derivative work of the training material. If people use it to create infringing (on the training or other material) they can be prosecuted in the exact same way they would if they created an infringing work via Photoshop or any other program. The same goes for other illegal uses such as creating harmful depictions of real people.

Accepting any expansion of IP rights, for whatever reason, would in fact be against the ethics of free software.

load more comments (5 replies)

load more comments (1 replies)

[–] lowleveldata 16 points 2 years ago (2 children)

trustworthy how?

[–] [email protected] 46 points 2 years ago* (last edited 2 years ago)

More transparent about data collection, and less likely to reinforce biases. Mozilla vision for trustworthy AI

[–] AdmiralShat 29 points 2 years ago (1 children)

Open source

[–] [email protected] 16 points 2 years ago (3 children)

I feel the issue with AI models isn’t their source not being open but the actual derived model itself not being transparent and auditable. The sheer number of ML experts who cannot explain how their model produces a given result is the biggest concern and requires a completely different approach to model architecture and visualisation to solve.

[–] [email protected] 21 points 2 years ago

Unfortunately, the nature of the models is that it's very difficult to get an understanding of the innards. That's part of the point, you don't need too. The best we can do is monitor how it's built and what connects in and out of it.

The open source bits let you see that it's not passing data on without your permission. If the training data is also open source, you can get for biases e.g. 90% of faces being white males.

[–] [email protected] 6 points 2 years ago

No amount of ML expertise will let someone know how a model produced a result, exactly. Training the model from the data requires a lot of very delicate math being done uncountable times to get a model that results in something useful, and it simply isn't possible to comprehend how the work inside is done in a meaningful way other than by doing guesswork.

load more comments (1 replies)

[–] [email protected] 14 points 2 years ago

Cautiously optimistic about this one.

[–] [email protected] 13 points 2 years ago* (last edited 2 years ago)

Couldn't give a fuck, there's already far too much bad blood regarding any form of AI for me.

It's been shoved in my face, phone and computer for some time now. The best AI is one that doesn't exist. AGI can suck my left nut too, don't fuckin care.

Give me livable wages or give me death, I care not for anything else at this point.

Edit: I care far more about this for privacy reasons than the benefits provided via the tech.

The fact these models reached "production ready" status so quickly is beyond concerning, I suspect the companies are hoping to harvest as much usable data as possible before being regulated into (best case) oblivion. It really no longer seems that I can learn my way out of this, as I've been doing since the beginning, as the technology is advancing too quickly for users, let alone regulators to keep it in check.

[–] [email protected] 10 points 2 years ago

Something that I'll definitely keep an eye on. Thanks for sharing!

[–] [email protected] 6 points 2 years ago

[–] [email protected] 5 points 2 years ago* (last edited 2 years ago) (1 children)

In which ways does this differ from stability ai, which made stable diffusion and also have a LLM afaik?

[–] [email protected] 10 points 2 years ago

The stability models aren't open source, the moralistic licence they release under violates the open source definition.

load more comments