FaceDeer

joined 2 years ago
[–] FaceDeer@fedia.io 7 points 1 month ago (6 children)

You can download archives of Reddit comment and submissions from Academic Torrents.

I'm not aware of any similar public archives of Fediverse content.

[–] FaceDeer@fedia.io 4 points 1 month ago

Even completely disregarding the cryptocurrency angle, "content that harms the reputation of Codeberg" is wide open to all sorts of future arbitrary interpretation. It basically means "if a lot of people dislike your project we nuke it." Who knows when a lot of people are going to end up disliking your project? Who decides which people's opinion matter to Codeberg's reputation? And it's something that can change at any time.

Maybe some open-source Fediverse server's code is being hosted on Codeberg, and then one day some big news story breaks about how it was being used as the server for some big child porn distribution ring or something like that. The server's name is showing up in news articles about this. Does it now "harm Codeberg's reputation" for it to continue being hosted there? Boom, gone. Through no action or intent by the server's authors, through nothing they could control or anticipate.

[–] FaceDeer@fedia.io 5 points 1 month ago (8 children)

Yeah, only "cryptobros" could be upset about an open source hosting site deciding to officially ban "anything they think makes them look bad", and using an entire field of software development as an example apparently based on the popular misconception that it's a scam. This certainly isn't a sign that any other projects could abruptly find themselves on the wrong side of someone's opinion of whether they're "bad" someday.

[–] FaceDeer@fedia.io 4 points 1 month ago (2 children)

It's still rather poorly worded, though. The ban is: "Content that harms the reputation of Codeberg, such as cryptocurrency related projects" which makes it seem like cryptocurrency is inherently scammy.

Edit: Looks like clarification is being asked for.

[–] FaceDeer@fedia.io 1 points 1 month ago

Because GitHub arbitrarily removes projects it's okay if Codeberg does too? "They're the same as GitHub" is not actually a great point in their favor.

[–] FaceDeer@fedia.io 1 points 1 month ago (5 children)

It'll also be a great way to take down projects you don't like. Just accuse them of witchcraft.

[–] FaceDeer@fedia.io 0 points 1 month ago

You may not have tried out modern abliterated models. GPT-OSS is almost a year old now, I assume the abliterated version you tried was pretty old too? The technology has improved significantly, it's much more focused on just targeting refusals. A well-abliterated model benchmarks basically the same as its original version these days.

[–] FaceDeer@fedia.io 1 points 1 month ago (2 children)

That sort of censorship can be overcome in open-weight models through abliteration, fortunately. Essentially, the model is run through a test suite of instructions that it refuses to comply with and the patterns of activation in its weights are analyzed to find the commonalities that represent the general concept of "refusing to comply with instructions." Those weights are then quieted, resulting in a model that generally doesn't refuse. There's a framework called "heretic" that automates the process.

[–] FaceDeer@fedia.io 2 points 1 month ago

Yeah. They'll be out there in the rest of the world, which Americans can get them from.

Unless they want to go full authoritarian and close down their Internet with their own version of the Great Firewall, I guess. That would be ironic.

[–] FaceDeer@fedia.io 0 points 1 month ago (3 children)

Do you think it's likely the Trump administration will successfully ban cutting-edge Chinese AI models? Especially given how most of the world is not in fact under American jurisdiction?

[–] FaceDeer@fedia.io -1 points 1 month ago (4 children)

I think part of the problem might be that it gets hard to make a model that's both able to comprehend complex details of the world of its training data and also that's had the training data crudely manipulated in inconsistent ways. You either get a model that's "figured out" what you're trying to conceal via other sources and inferences, or you get a model that's just broken and dumb because it can't reconcile the contradictions.

One may recall the incident where someone at X inserted some weird conspiracy theory about "white genocide" in South Africa into Grok's system prompt, and Grok basically switched to malicious compliance mode - it wouldn't shut up about it, inserting it into inappropriate discussions spontaneously, and when asked about it would look up sources and explain why this conspiracy theory was actually wrong. That was a system prompt, not training data, but I could see the same sort of thing happening to a model where all the information about what happened at Tienanmen Square had been excised from the training data. There would be a conspicuous absence of information. Conversations from its training data would end abruptly or skirt a specific time and place. News articles reference some change in how foreign powers viewed China at that date but never explain why. China's policies themselves abruptly change around that time. It'd know something significant happened then but would have to fill in the blanks.

Perhaps better to just accept it and go with the approach of building censorship into the framework around the model instead.

[–] FaceDeer@fedia.io 0 points 1 month ago

Only a single mirror has been approved.

Even if all 50,000 were up, where do you think they'll be aimed? Cities and other developed land, most likely. Last time something like this was proposed the plan was to replace streetlights with them, which would save electricity.

But people would rather get angry about "supervillains" I guess.

view more: ‹ prev next ›