AI & AUTOMATION

Autoresearch by Andrej Karpathy: Autonomous AI Research Agents for Single-GPU Nanochat Training and LLM Training Optimization

Autoresearch shifts from manual hyperparameter tuning to agent-led discovery, where autonomous LLM agents iteratively formulate hypotheses, modify code, execute fixed-time nanochat training runs, and analyze results to uncover optimal configurations overnight.

Introduction

Autoresearch represents a foundational advance in agentic research, a paradigm where large language model agents autonomously conduct scientific inquiry within a constrained but realistic machine learning environment. Developed by Andrej Karpathy, this open-source tool equips LLM agents with a minimal single-GPU nanochat training setup, enabling them to explore architectures, optimizers, and hyperparameters without human intervention. Instead of researchers manually iterating through code edits and loss curve inspections, the agent handles the full cycle of discovery, turning overnight compute into a stream of validated experiments.

The framework builds directly on a simplified implementation of nanochat, Karpathy’s lightweight GPT-style model. By constraining each training run to a fixed 5-minute wall-clock budget, Autoresearch ensures every experiment remains comparable regardless of hardware variations. The sole evaluation metric—validation bits per byte (val_bpb)—rewards genuine improvements in model efficiency. This design makes Autoresearch an ideal entry point for machine learning researchers and software engineers seeking automated machine learning agents that accelerate LLM training optimization while preserving interpretability through git-tracked changes.

Core Workflow

The agentic research loop in Autoresearch follows a disciplined, repeatable sequence that mirrors traditional scientific methodology yet executes entirely under agent control. The process begins with the agent reviewing the current git state and the human-provided instructions in program.md. It then formulates a hypothesis—such as altering model depth, switching optimizers, or adjusting batch sizes—and directly edits the single modifiable file, train.py.

After committing the change on a dedicated branch (e.g., autoresearch/mar18), the agent executes uv run train.py. Training halts precisely after 5 minutes of wall-clock compute time, excluding startup and compilation overhead. The agent parses the output log for the critical val_bpb value and peak VRAM usage, deciding whether to keep the commit (if val_bpb improved) or revert via git reset (if results worsened or stayed equal). Results are logged to an untracked results.tsv file with commit hash, metric, memory, status (“keep”, “discard”, or “crash”), and a concise description.

This loop repeats indefinitely—approximately 12 experiments per hour—until manually interrupted. The fixed time budget and single-file modification policy guarantee that every iteration remains auditable and comparable, transforming nanochat training into a self-improving system where agents discover superior configurations through relentless, data-driven iteration.

Environment Setup

Setting up Autoresearch requires a single NVIDIA GPU and a minimal Python environment managed by the modern project tool uv. The process is deliberately lightweight to support rapid experimentation on accessible hardware.

Begin by installing uv if it is not already present:

curl -LsSf https://astral.sh/uv/install.sh | sh

Clone the repository and navigate into the directory:

git clone https://github.com/karpathy/autoresearch.git
cd autoresearch

Install dependencies and activate the managed environment:

uv sync

Perform one-time data preparation, which downloads shards and trains the BPE tokenizer:

uv run prepare.py

Verify the baseline setup by running a manual training pass:

uv run train.py

This command should complete in approximately 5 minutes and output a summary including val_bpb, peak VRAM, and model statistics. No OpenAI API key or external LLM server is required; the autonomous agent operates through direct prompting of an external chat interface such as Claude or Codex. Ensure CUDA is available and that your GPU meets the soft VRAM constraints outlined in program.md.

Practical Usage

Launching an autonomous research session begins with configuring the research goal inside program.md—the human-editable Markdown file that serves as the agent’s complete instruction set and skill definition. Edit this file to refine objectives, constraints, or success criteria before starting any run.

Create a fresh experiment branch using a date-based tag:

git checkout -b autoresearch/mar18

Initialize results.tsv with the header row only. Then open your preferred LLM chat interface (Claude recommended) with the repository context loaded and issue the starter prompt:

Hi have a look at program.md and let's kick off a new experiment! let's do the setup first.

The agent will handle branch setup, baseline run, and subsequent iterations without further input. Monitor progress by inspecting the git log, the accumulating results.tsv (kept untracked), and the run.log files generated per experiment. To review overall trends, open the included analysis.ipynb notebook.

For a concrete example of the Research Plan template, the repository’s default program.md provides the exact structure agents follow. Key sections include:

## Setup
To set up a new experiment, work with the user to:
1. Agree on a run tag...
2. Create the branch...
3. Read the in-scope files...

## Experimentation
Each experiment runs on a single GPU. The training script runs for a fixed time budget of 5 minutes...

## The experiment loop
LOOP FOREVER:
1. Look at the git state...
2. Tune train.py...
3. git commit...
4. Run the experiment...
5. Read out the results...
6. Record the results in the tsv...
7. If val_bpb improved... keep; else git reset...

This template can be extended with additional agents, complexity penalties, or domain-specific hypotheses, enabling sophisticated automated machine learning agents tailored to specific nanochat training goals.

Technical Analysis

Autoresearch matters profoundly for the future of AI because it demonstrates the first practical step toward self-improving systems at the research level. Rather than relying on human researchers to propose and validate every architectural tweak, the framework delegates hypothesis generation, code modification, execution, and statistical analysis to LLM agents. Over the course of a night, a single GPU can produce 100+ comparable experiments—far exceeding what most individuals achieve manually—while the git history and results.tsv provide complete provenance for every decision.

The deliberate focus on single-GPU setups and fixed-time budgets addresses a critical efficiency bottleneck in LLM training optimization. Traditional hyperparameter sweeps require massive compute clusters and weeks of orchestration; Autoresearch collapses that complexity into an overnight process that remains fully reproducible and auditable. By targeting nanochat training—a minimal yet realistic GPT implementation—the tool democratizes frontier experimentation for researchers without access to supercomputers.

In essence, Autoresearch inaugurates the era of autonomous AI research. It proves that automated machine learning agents can systematically outperform manual exploration in constrained environments, paving the way for larger-scale swarms that will eventually drive the next generation of model architectures and training paradigms. Researchers and engineers are encouraged to fork the repository, customize program.md, and contribute notable platform adaptations, further accelerating collective progress in open-source LLM training optimization.

You may also like

Subscribe
Notify of
guest

0 Comments
Newest
Oldest Most Voted