Get Started
Topic · 100 recaps

AI Agents

Research on AI systems that plan, use tools, and act autonomously across multi-step tasks — from chat-only assistants to embodied robotic agents.
PostsQuestions
Home
Topics
Search
Library
Sort
Newest
$Φ$-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them?
Agents · Sep 9
Show-Harness: Just a VLM Agent Can Play Robots
Agents · Sep 9 · 10:04
Agent Memory Controls Must Follow Consequences, Not Labels
Agents · Sep 8 · 10:35
SchemeArena: Factorized Stress Testing of Scheming in LLM Agents
Agents · Sep 8
AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems
Agents · Sep 8
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness
Agents · Sep 8
Procedural Graphs: Self-Evolving Execution Structures for LLM Agents
Agents · Sep 8
Ask Before You Optimize: Dynamic Pre-Formulation Clarification for Interactive Optimization
Agents · Sep 8
Three Ways "Our Agent's Memory Still Works" Can Be Wrong
Agents · Sep 7 · 12:45
Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents
Agents · Sep 7