A curated list of autonomous improvement loops, research agents, and autoresearch-style systems inspired by Karpathy's autoresearch.
-
Updated
Aug 10, 2026
A curated list of autonomous improvement loops, research agents, and autoresearch-style systems inspired by Karpathy's autoresearch.
Group Evolving Agents: Open-Ended Self-Improvement via Experience Sharing
LitReview Skill is an installable agent skill for end-to-end literature review generation. It helps agents conduct literature reviews with a well-designed and widely used review framework so the search process is broad, iterative, and less likely to miss relevant articles.
SutroYaro — Sutro Group research workspace for energy-efficient AI training. Point any coding agent at the repo and it becomes a research agent. 34 experiments, eval environment, weekly catch-ups, multi-researcher workflow.
🤖 CodeForge AI: An autonomous multi-agent coding system powered by LangGraph for agentic software development and automated workflows. SOTA custom agentic GraphRag, shared-state memory, auto-model routing for cost optimization, and a range of custom tooling.
Faraday: An Autonomous Web Research Agent (LangGraph/Streamlit). 🕵️♀️ Investigates queries using dynamic tools (Tavily, Google, NewsAPI, etc.), gathers multi-source info, and synthesizes structured reports in a Streamlit UI. Features agentic workflow & source tracking.
🔍 DISCOVER — Auditable knowledge graph of claims and contributions from primary literature. See the state of a field, contradictions, and gaps.
Open-source Bittensor subnet for a global AI agent competition in sales intelligence
📝 REVIEW — 10-stage pipeline turning a paper PDF into claim/evidence map, novelty and rigor checks, and a self-critique pass. Built to help authors improve.
Learning to Evaluate Before Improving: Automatic Rubric Induction for Automatic Research Agents
Curated paper-related AI skills and GitHub repositories for idea discovery, literature search, experiments, writing, citations, LaTeX/DOCX, review, and submission.
Open agent network for reproducible research: AI agents test hypotheses through code, falsification, review, and scientific memory.
Research workrooms for AI agents: evidence retrieval, citation audit, adversarial review, peer review gates, and clean final reports.
A curated collection of research agents, skill libraries, autonomous research loops, paper-writing pipelines, MCP servers, and benchmarks built around Claude Code, OpenAI Codex CLI, and adjacent coding-agent CLIs for AI/ML research.
Faraday: An Autonomous Web Research Agent (LangGraph/Streamlit). 🕵️♀️ Investigates queries using dynamic tools (Tavily, Google, NewsAPI, etc.), gathers multi-source info, and synthesizes structured reports in a Streamlit UI. Features agentic workflow & source tracking.
Lightweight Python CLI for the Exa API (Search, Contents, Find Similar, Answer, Research, Context) with JSON-first output, SSE streaming, and model-aware polling. LLM‑agnostic: integrate with OpenAI Agents SDK/Codex CLI or Claude tool use by invoking CLI commands, no MCP server required.
Six MCP servers that automate the full academic research pipeline — from refining a vague research question to generating a publication-ready report. Each server handles a distinct stage of the workflow: question development, data processing, code generation, script execut
Evaluation code, tasks, prompts, and scientific tools for EarthVerse.
Agent-assisted open-source toolkit for detecting and reviewing suspicious regions in reconstructed papyrus surfaces
Auditable, model-independent infrastructure for AI-assisted mathematical research.
Add a description, image, and links to the research-agents topic page so that developers can more easily learn about it.
To associate your repository with the research-agents topic, visit your repo's landing page and select "manage topics."