Rankings

Top AI Research Papers

The overall board: every AI paper from the last 90 days, across language models, agents, vision, robotics and the rest. Ranked by three independent signals — GitHub star velocity on the official implementation, citation count, and Hugging Face attention — so a paper that only trended for a day does not hold a place. Rebuilt every day.

Ranking window: 2026-07-04 – 2026-10-01 Rankings last rebuilt 2026-10-02

3

Raven: an open ecosystem that builds, evolves and orchestrates agent harnesses for long-horizon tasks

Raven argues that hand-crafted agent harnesses are too domain-coupled to scale with long-horizon, cross-domain workloads. It treats each model–harness pair as a composable unit of intelligence: a Host Agent decomposes goals and coordinates specialized sub-agents, a host archive and EverOS retain experience across tasks, and Skill Forge turns that experience into reusable procedures — with theory on expanding reliable task coverage under one resource budget.

  • 1,007 stars/7d
  • 0 citations
  • 493 upvotes
4

LingBot-World 2.0: a game world model with an unbounded interaction horizon and a 60 fps real-time variant

LingBot-World 2.0 upgrades the game world model to an unbounded interaction horizon with consistent output quality via a causal pretraining paradigm, distills a real-time variant fast enough to drive 720p video at 60 fps, and greatly expands the action set (attacking, archery, spell-casting, shooting) plus text-driven interactive elements.

  • 22 stars/7d
  • 17 citations
  • 47 upvotes
7

LingBot-VLA 2.0: a vision-language-action model pretrained on 60,000 hours spanning 20 robot embodiments

This technical report upgrades LingBot-VLA with a revamped data pipeline of about 60,000 pretraining hours: 50K robot-trajectory hours across 20 embodiments plus 10K egocentric human-video hours. It extends control beyond dual arms to heads, waists, mobile bases and dexterous hands, aiming squarely at the gap between lab robotics and real deployment.

  • 32 stars/7d
  • 15 citations
  • 21 upvotes
  • #3 robotwin-2-0-easy-50-tasks
8

RRSI: regularizing recursive self-improvement so evolved agent harnesses survive out-of-distribution benchmarks

Iteratively editing an agent's harness (prompts, control flow, tooling, memory) can look like self-improvement while quietly overfitting the training tasks. RRSI regularizes both ends of the loop: the proposer works under a temporally annealed budget and is steered toward unexplored trajectories, while the selector uses a critic to screen benchmark-specific proposals and a pruner to drop edits that are too small, too expensive, or no longer useful. The authors report gains of up to 14.1 points on the evolved split and up to 4.7 points on five held-out benchmarks, with a harness that uses 30% fewer policy tokens.

  • 843 stars/7d
  • 0 citations
  • 172 upvotes
9

Xiaomi Robotics 1: a vision-language-action model pretrained on over 100K hours of real-world trajectories

Xiaomi's robotics team presents a vision-language-action foundation model for mobile manipulation, pretrained on more than 100K hours of real-world UMI trajectories with auto-labeled scene-transition language, then post-trained to match embodiments and human imperative prompts. The paper claims clean scaling with more data and parameters, out-of-the-box performance in unseen environments and efficient fine-tuning for dexterous tasks.

  • 9 stars/7d
  • 15 citations
  • 75 upvotes
  • #3 robocasa
11

RoboDojo: one benchmark that puts generalist robot policies through 42 simulation and 18 real-world tasks

RoboDojo is a unified sim-and-real benchmark for generalist robot manipulation policies, with 42 simulation tasks and 18 real-world tasks probing generalization, memory, precision and long-horizon control. It ships parallel Isaac Sim evaluation plus a cloud-accessible real-eval system and a leaderboard the authors populated with 30 policies.

  • 35 stars/7d
  • 12 citations
  • 17 upvotes
12

YuE2: one Mixture-of-Transformers that writes a readable score first, then renders full-song audio from it

Symbolic music models plan composition explicitly but never sound like a finished recording, while audio models produce full songs that keep the composition hidden. YuE2 combines both in a single AR-NAR Mixture-of-Transformers that plans symbolically, expands the plan into semantic music tokens, and renders the full song, using two new representation models (MERT2, SheetSage2) to learn from recordings that lack aligned scores. The authors report a SongBench Global Avg of 6.73 on WildSongBench (6.96 at best-of-8), and in expert listening the best-of-8 samples were preferred over Suno v4.5 and roughly tied with Suno v5; the readable score also makes the output editable by external language-model agents.

  • 437 stars/7d
  • 0 citations
  • 216 upvotes
13

WROP: a cognitive-science exam that tests whether video world models understand object permanence — and a 16B model trained to pass it

When an object leaves the frame, does a video generation model remember it exists? The authors built WROP, 150 cognitive-science tasks across six categories rendered in Blender with randomized lighting, speed, and camera angle — a 1.5M-sample training corpus plus a 300-question exam used to test 14 video models. They then trained PWM-WROP, a 16B world model on their native PyTorch stack for AWS Trainium2, releasing data, exam, scores, and weights. In a blind pairwise Elo study, it ranked first among continuation models and effectively tied the best reference-to-video models.

  • 409 stars/7d
  • 0 citations
  • 193 upvotes
15

LLM-as-a-Verifier: continuous verification scores read from scoring-token logits, no judge prompts needed

LLM-as-a-Verifier turns a language model into a verifier by computing expectations over its scoring-token logits instead of prompting for discrete judgments, yielding continuous feedback with no extra training. The authors report state-of-the-art results on benchmarks including Terminal-Bench V2 and SWE-Bench Verified, and ship a Claude Code extension that watches agents and supplies dense rewards for RL.

  • 24 stars/7d
  • 11 citations
  • 19 upvotes
16

A survey of self-improving agents: how a foundation model plus scaffold turns experience into capability

This survey decomposes a self-improving agent into a foundation model and an operational scaffold of prompts, memory, tools and control logic, then defines improvement as a self-induced update operator that converts experience into accumulated capability with minimal human input. It organizes the growing literature by update target and driving signal, from prompt and memory edits to weight-level self-modification.

  • 19 stars/7d
  • 9 citations
  • 35 upvotes
17

Frontis-MA1: an open 35B meta-evolution agent for recursive self-improvement in ML engineering

Frontis-MA1 ships with OpenMLE, an open full-stack system (task gyms, RL operator learning, evolutionary search) for studying recursive self-improvement. Its 35B meta-evolution agent, post-trained around Draft, Improve, Debug and Crossover program-evolution operators, claims MLE-Bench Lite results the authors say exceed GPT-5.5 plus Codex, with components transferring to held-out NatureBench Lite.

  • 26 stars/7d
  • 3 citations
  • 186 upvotes
18

ABot-World-0: an action-conditioned world model streaming interactive 720p worlds on a single desktop GPU

ABot-World-0 is an action-conditioned video world model trained on multi-source data from AAA games, simulation engines and internet videos, controlled through raw keyboard input with reference-character memory for consistent third-person rollouts. The authors report streaming 720p at up to 16 FPS on one RTX 5090 with 1.2s action-to-first-frame latency, targeting real-time long-horizon closed-loop interaction.

  • 6 stars/7d
  • 6 citations
  • 313 upvotes
19

Dream-RSI: coding agents that get better at exploring by replaying their own past discoveries offline

Dream-RSI is a framework for letting coding agents improve their own exploration strategy instead of keeping a fixed one. It builds a replay simulator from the agent's past discovery trees, lets the exploration policy practice inside that simulator for cheap off-policy feedback, then redeploys the refined strategy online so new discoveries feed the next round. The authors report competitive or better discovery quality on algorithm engineering, math optimization, and GPU kernel tasks, with substantially lower discovery cost in several settings.

  • 173 stars/7d
  • 0 citations
  • 270 upvotes

How this list is ranked

The overall board: every AI paper from the last 90 days, across language models, agents, vision, robotics and the rest. Ranked by three independent signals — GitHub star velocity on the official implementation, citation count, and Hugging Face attention — so a paper that only trended for a day does not hold a place. Rebuilt every day.

A dash means the signal was unavailable, not zero.