Rankings

Best AI Agent Papers

Agent papers from the last 90 days — tool use, computer use, coding agents, multi-agent systems. Ranked by three independent signals: GitHub star velocity on the official implementation, citation count, and Hugging Face attention. Papers strong on only one signal are left out. Rebuilt every day.

Ranking window: 2026-07-04 – 2026-10-01 Rankings last rebuilt 2026-10-02

2

LingBot-World 2.0: a game world model with an unbounded interaction horizon and a 60 fps real-time variant

LingBot-World 2.0 upgrades the game world model to an unbounded interaction horizon with consistent output quality via a causal pretraining paradigm, distills a real-time variant fast enough to drive 720p video at 60 fps, and greatly expands the action set (attacking, archery, spell-casting, shooting) plus text-driven interactive elements.

  • 22 stars/7d
  • 17 citations
  • 47 upvotes
3

Raven: an open ecosystem that builds, evolves and orchestrates agent harnesses for long-horizon tasks

Raven argues that hand-crafted agent harnesses are too domain-coupled to scale with long-horizon, cross-domain workloads. It treats each model–harness pair as a composable unit of intelligence: a Host Agent decomposes goals and coordinates specialized sub-agents, a host archive and EverOS retain experience across tasks, and Skill Forge turns that experience into reusable procedures — with theory on expanding reliable task coverage under one resource budget.

  • 1,007 stars/7d
  • 0 citations
  • 493 upvotes
5

RRSI: regularizing recursive self-improvement so evolved agent harnesses survive out-of-distribution benchmarks

Iteratively editing an agent's harness (prompts, control flow, tooling, memory) can look like self-improvement while quietly overfitting the training tasks. RRSI regularizes both ends of the loop: the proposer works under a temporally annealed budget and is steered toward unexplored trajectories, while the selector uses a critic to screen benchmark-specific proposals and a pruner to drop edits that are too small, too expensive, or no longer useful. The authors report gains of up to 14.1 points on the evolved split and up to 4.7 points on five held-out benchmarks, with a harness that uses 30% fewer policy tokens.

  • 843 stars/7d
  • 0 citations
  • 172 upvotes
6

LLM-as-a-Verifier: continuous verification scores read from scoring-token logits, no judge prompts needed

LLM-as-a-Verifier turns a language model into a verifier by computing expectations over its scoring-token logits instead of prompting for discrete judgments, yielding continuous feedback with no extra training. The authors report state-of-the-art results on benchmarks including Terminal-Bench V2 and SWE-Bench Verified, and ship a Claude Code extension that watches agents and supplies dense rewards for RL.

  • 24 stars/7d
  • 11 citations
  • 19 upvotes
7

A survey of self-improving agents: how a foundation model plus scaffold turns experience into capability

This survey decomposes a self-improving agent into a foundation model and an operational scaffold of prompts, memory, tools and control logic, then defines improvement as a self-induced update operator that converts experience into accumulated capability with minimal human input. It organizes the growing literature by update target and driving signal, from prompt and memory edits to weight-level self-modification.

  • 19 stars/7d
  • 9 citations
  • 35 upvotes
9

ABot-World-0: an action-conditioned world model streaming interactive 720p worlds on a single desktop GPU

ABot-World-0 is an action-conditioned video world model trained on multi-source data from AAA games, simulation engines and internet videos, controlled through raw keyboard input with reference-character memory for consistent third-person rollouts. The authors report streaming 720p at up to 16 FPS on one RTX 5090 with 1.2s action-to-first-frame latency, targeting real-time long-horizon closed-loop interaction.

  • 6 stars/7d
  • 6 citations
  • 313 upvotes
10

Dream-RSI: coding agents that get better at exploring by replaying their own past discoveries offline

Dream-RSI is a framework for letting coding agents improve their own exploration strategy instead of keeping a fixed one. It builds a replay simulator from the agent's past discovery trees, lets the exploration policy practice inside that simulator for cheap off-policy feedback, then redeploys the refined strategy online so new discoveries feed the next round. The authors report competitive or better discovery quality on algorithm engineering, math optimization, and GPU kernel tasks, with substantially lower discovery cost in several settings.

  • 173 stars/7d
  • 0 citations
  • 270 upvotes
13

Ouroboros: a coding agent that evolves its own tools, prompts and core through reviewed commits

Ouroboros is a self-developing coding-agent harness whose tools, prompts, context assembly and even core implementation improve through reviewed commits that become its later runtime, with two evolution modes: recursive free evolution and experience-driven core evolution. The authors also report a 161-day live deployment across seven human communication surfaces, arguing safety guardrails must stay authoritative while the agent rewrites its own code.

  • 25 stars/7d
  • 2 citations
  • 92 upvotes
  • #1 osworld-verified
14

EvoOntology: a self-evolving semantic layer that lets data agents query heterogeneous data through an MCP server

EvoOntology tackles the agent-data gap: agents can see column names and file paths through generic tools, but not what the data actually means. It wraps an ontology (schema, content, and tool layers) as an MCP server the agent can query at runtime, uses a builder agent to construct it autonomously, and keeps refining it with attribution-guided edits that are only accepted after a paired evaluation against the backbone. The authors report consistent gains over strong baselines and hand-built semantic layers across three data-agent benchmarks and four LLM backbones.

  • 164 stars/7d
  • 0 citations
  • 126 upvotes

How this list is ranked

Agent papers from the last 90 days — tool use, computer use, coding agents, multi-agent systems. Ranked by three independent signals: GitHub star velocity on the official implementation, citation count, and Hugging Face attention. Papers strong on only one signal are left out. Rebuilt every day.

A dash means the signal was unavailable, not zero.