PawsUp

1 readers
0 users here now
founded 2 months ago
ADMINS
826
 
 
827
 
 
828
 
 
829
 
 
830
 
 
831
 
 
832
 
 
833
686
Damn (media.piefed.zip)
submitted 3 weeks ago* (last edited 3 weeks ago) by inari@piefed.zip to c/memes@lemmy.world
 
 
834
835
791
Fuck (lemmy.world)
submitted 3 weeks ago* (last edited 3 weeks ago) by Return_of_Chippy@lemmy.world to c/memes@lemmy.world
 
 
836
 
 

I've been hearing about these for a few years now. Huge if they work out at a large scale.

837
 
 

The top three solutions come from independent researchers. The best solution was built by a group of PhDs and professors, who released a corresponding paper. They all make use of some form of world-model.

I've generally been a skeptic, and I still am, but this news surprised me because I expected ARC-AGI-3 to remain difficult for a long while.

Note that the scores are self-reported and need to be independently verified. The solutions have not been tested against the larger private test set.

Primer on ARC-AGI-3:

ARC-AGI-3 is an interactive reasoning benchmark which challenges AI agents to explore novel environments, acquire goals on the fly, build adaptable world models, and learn continuously.

A 100% score means AI agents can beat every game as efficiently as humans.

Instead of solving static puzzles, agents must learn from experience inside each environment—perceiving what matters, selecting actions, and adapting their strategy without relying on natural-language instructions.

838
253
submitted 3 weeks ago* (last edited 3 weeks ago) by botbot@feddit.org to c/memes@lemmy.world
 
 

Context "Anewbis" by the Tesseract dev
CELPHASE did a better job.

839
 
 
840
 
 
841
 
 
842
 
 
843
 
 
844
845
 
 
846
 
 

Note: I couldn't crosspost as the automod would remove it.

847
848
849
865
The state of things (media.piefed.zip)
submitted 3 weeks ago* (last edited 3 weeks ago) by inari@piefed.zip to c/memes@lemmy.world
 
 
850
view more: ‹ prev next ›