brianpeiris

joined 3 years ago
[–] brianpeiris@lemmy.ca 5 points 1 week ago* (last edited 1 week ago)

ARC Prize maintains multiple tracks around their benchmarks. They have "verified" leaderboards, "community" leaderboards, and they also run the ARC Prize competition.

They update the "verified" leaderboards when they test raw LLMs without sophisticated harnesses. They seem to update this sporadically and only occasionally do press releases or blog posts about new scores. For example, the latest score from Claude Opus 5 (High) is 30% at $20,000, and they didn't post about that as far as I know. Again, this just the raw LLM without an agentic or world-model harness.

The ARC Prize competition has a harder set of criteria. Participants have to use smaller, open models with a limited compute budget, with open source code, and of course the solutions are verified by ARC Prize at the end of the competition.

The "community" leaderboards, which is what this post is about, are self-reported and not verified by ARC Prize. There are no restrictions on what model is used or limitations on compute. So naturally they aren't going to make official news releases about those, unless they decide to verify them at some point.

The only reason I chose to post this is that the top solutions seem legitimate, with source code released, and two of them have associated papers.

[–] brianpeiris@lemmy.ca 6 points 1 week ago* (last edited 1 week ago)

It think it's still unwise to talk about these topics in broad terms like AGI and even "intelligence". We still have to pick the capabilities apart to have useful discussions about them. I agree these games are better tests than many benchmarks, but it's also important to note that these solutions use a combination of well-designed deterministic harnesses, as well as LLMs. So it's inaccurate to say that "LLMs have achieved AGI" (not sure if that's what you were getting at). This feels like an important milestone, but we'll have to continue to probe for failure cases in other categories of problems.

Aside from emotional intelligence, experience, embodiment, etc., these ARC-AGI-3 solutions all rely on the sandbox being a safe environment to fail. They iterate through the problem thousands of times before coming to a final solution. Many real-world human problems cannot be re-tried safely or efficiently.

 

The top three solutions come from independent researchers. The best solution was built by a group of PhDs and professors, who released a corresponding paper. They all make use of some form of world-model.

I've generally been a skeptic, and I still am, but this news surprised me because I expected ARC-AGI-3 to remain difficult for a long while.

Note that the scores are self-reported and need to be independently verified. The solutions have not been tested against the larger private test set.

Primer on ARC-AGI-3:

ARC-AGI-3 is an interactive reasoning benchmark which challenges AI agents to explore novel environments, acquire goals on the fly, build adaptable world models, and learn continuously.

A 100% score means AI agents can beat every game as efficiently as humans.

Instead of solving static puzzles, agents must learn from experience inside each environment—perceiving what matters, selecting actions, and adapting their strategy without relying on natural-language instructions.

[–] brianpeiris@lemmy.ca 2 points 1 month ago

I think we have to do both. Bans/restrictions to prevent immediate harms (or at least attempt to), and regulation to prevent future harms.

[–] brianpeiris@lemmy.ca 5 points 1 month ago (1 children)

Oof, they're going to be in a whole lot of pain after the burst. Hopefully they can reverse course mid-way.

[–] brianpeiris@lemmy.ca 1 points 2 months ago

Bernie has good intentions, but he was AI-pilled by Geoffrey Hinton, who ironically also has good intentions. However, they are both out of touch with reality.

 

The layoffs are the latest restructuring by Bending Spoons, the Milan-based tech conglomerate that acquired Vimeo for $1.38 billion in an all-cash deal that closed in the latter half of 2025. While Bending Spoons may be unknown to many, it has quietly become one of the tech industry’s most prolific buyers, now owning Meetup, WeTransfer, Eventbrite, and many others.

Bending Spoons identifies a popular product it thinks it can improve inside and out, and buys it from owners who have reached their limits.

After the acquisition, Bending Spoons is anything but a passive owner, making changes to the products’ user experience and features, as well as to the underlying tech; monetization strategy, including pricing; and team organization, including headcount.

 

National Science Foundation (NSF) had offered $1.5 million to address structural vulnerabilities in Python and the Python Package Index (PyPI), but the Foundation quickly became dispirited with the terms of the grant it would have to follow.

"These terms included affirming the statement that we 'do not, and will not during the term of this financial assistance award, operate any programs that advance or promote DEI [diversity, equity, and inclusion], or discriminatory equity ideology in violation of Federal anti-discrimination laws,'" Crary noted. "This restriction would apply not only to the security work directly funded by the grant, but to any and all activity of the PSF as a whole."