Nanochat by Karpathy: The Simplest Experimental Harness for Training LLMs and Building Your Own ChatGPT on a Single GPU
nanochat delivers the best ChatGPT-level model that $100 can buy by offering the simplest, single-dial compute-optimal harness for full-pipeline LLM training and inference on accessible hardware.
Introduction
nanochat stands as the definitive minimal experimental harness for training large language models, created by Andrej Karpathy. Described as “the best ChatGPT that $100 can buy,” it distills the entire LLM pipeline—tokenization, pre-training, supervised fine-tuning, reinforcement learning, evaluation, inference, and a production-grade web UI—into a single, hackable codebase that runs efficiently on one GPU node.
The philosophy is uncompromising simplicity: no sprawling configuration factories, no conditional complexity, and only one user-controlled parameter—the model depth. All other hyperparameters (width, number of heads, learning-rate schedule, training horizon, weight decay, batch size) are derived automatically to maintain compute optimality at every scale. This design makes nanochat ideal for researchers, engineers, and enthusiasts who want to iterate rapidly on LLM architectures without the overhead of production-scale frameworks.
By targeting a GPT-2-capability model (approximately 26 layers), nanochat achieves competitive performance in roughly two hours on an 8xH100 node for about $48, or $15 on spot instances. The framework builds directly on lessons from nanoGPT while extending the pipeline to modern chat-model workflows, including a fully functional ChatGPT-style interface. Its focus on single-node execution and automatic hyperparameter scaling democratizes frontier LLM experimentation, positioning nanochat as the essential open-source tool for self-hosted LLM training.

Core Features and Compute-Optimal Design
At the heart of nanochat lies a deliberate minimalism that prioritizes hackability and reproducibility. The core GPT Transformer implementation resides in a single cohesive file, with supporting modules for tokenization, data loading, optimization, and evaluation kept equally concise. Precision is managed globally through the NANOCHAT_DTYPE environment variable, defaulting to bfloat16 on modern GPUs and float32 elsewhere, eliminating the need for manual autocast wrappers.
The sole complexity dial is --depth, which controls the number of transformer layers. Everything else—model width, attention heads, learning rate schedule, and total training steps—scales automatically according to established compute-optimal laws. This approach eliminates months of manual hyperparameter search, allowing users to focus on architectural innovations rather than tuning.
Evaluation integrates two rigorous metrics: the DCLM CORE score (targeting GPT-2’s 0.256525 benchmark) and bits-per-byte validation. A public “Time-to-GPT-2” leaderboard tracks wall-clock efficiency, fostering community-driven optimization. The harness further supports full fine-tuning pathways, including supervised fine-tuning (SFT) and reinforcement learning scripts, plus synthetic data generation for identity tuning.
Installation Guide
Setting up nanochat requires only a modern GPU environment and the lightweight uv package manager. The process is intentionally streamlined to enable immediate experimentation.
Begin by cloning the repository:
git clone https://github.com/karpathy/nanochat.git
cd nanochatCreate and activate a virtual environment:
uv venv .venv
source .venv/bin/activateInstall the project in editable mode to pull all dependencies defined in pyproject.toml:
uv pip install -e .The installation automatically resolves PyTorch and other requirements. No additional CUDA setup beyond a standard driver is needed; the framework auto-detects hardware capabilities. For CPU or Apple Silicon testing, the same steps apply, though performance will be significantly reduced.
Verify the environment by checking the optional dtype variable:
echo $NANOCHAT_DTYPELeave it unset for automatic detection (bfloat16 on H100-class GPUs).
Data Preparation and Prerequisites
nanochat ships with reference scripts for data handling but assumes users supply their own pre-training corpus. The recommended approach uses high-quality datasets such as NVIDIA ClimbMix or similar public mixtures. Run the provided reference pipeline:
python dev/repackage_data_reference.pyThis script converts raw text into tokenized shards compatible with the distributed data loader. For smaller experiments, the repository includes lightweight synthetic data generators:
python dev/gen_synthetic_data.pyHardware prerequisites emphasize accessibility: an 8xH100 node is optimal, but the code supports any single GPU through gradient accumulation by reducing --device_batch_size (default 32). Minimum VRAM is approximately 80 GB for full-batch runs; lower-memory cards simply require smaller device batches. The framework requires CUDA compute capability 8.0 or higher for bfloat16 acceleration but gracefully falls back on older hardware.
Training Workflows
Training begins with the most straightforward path: the speedrun script for a full GPT-2-level model.
Execute the complete training sequence:
bash runs/speedrun.shThis command launches an 8xH100-optimized run using torchrun, completing in approximately three hours and producing a model exceeding the GPT-2 CORE benchmark. Upon completion, launch the interactive web UI:
python -m scripts.chat_webAccess the ChatGPT-like interface at http://<your-node-ip>:8000. The UI connects directly to the trained checkpoint for real-time inference.
For rapid iteration and code experimentation, use the minimal training command:
OMP_NUM_THREADS=1 torchrun --standalone --nproc_per_node=8 -m scripts.base_train -- \
--depth=12 \
--run="d12" \
--model-tag="d12" \
--core-metric-every=999999 \
--sample-every=-1 \
--save-every=-1This five-minute experiment trains a depth-12 model with Weights & Biases logging under the specified run name. Adjust --depth freely to explore scaling behavior.
Additional specialized runs include:
bash runs/scaling_laws.sh # systematic depth series
bash runs/miniseries.sh # multiple compute-optimal models
bash runs/runcpu.sh # CPU-only verificationAll training scripts integrate seamlessly with wandb for experiment tracking and produce checkpoints compatible with the inference engine.
Inference, Evaluation, and Chat Interfaces
nanochat provides both command-line and web-based interaction out of the box. For quick testing:
NANOCHAT_DTYPE=bfloat16 python -m scripts.chat_cli -p "Explain compute-optimal scaling in LLMs"The CLI supports streaming responses with KV-cache acceleration for low latency.
For evaluation, dedicated scripts compute rigorous metrics:
python -m scripts.core_eval
python -m scripts.loss_evalThese assess CORE score and bits-per-byte on held-out validation sets. Fine-tuning pathways are equally accessible:
python -m scripts.chat_sft # supervised fine-tuning
python -m scripts.chat_rl # reinforcement learning stageThe web UI (scripts/chat_web) serves a clean, browser-based chat interface that mirrors commercial systems while remaining fully self-hosted and open source.
Advanced Capabilities and Research Applications
Beyond basic training, nanochat supports tokenizer customization through scripts/tok_train.py and compression evaluation via scripts/tok_eval.py. Researchers can inject custom tasks from the tasks/ directory to benchmark against ARC, GSM8K, MMLU, and other standards.
The automatic hyperparameter derivation makes nanochat particularly powerful for scaling-laws studies. By varying only --depth and observing resulting performance, users can replicate and extend classical compute-optimal research without manual tuning overhead. The modular engine (nanochat/engine.py) and optimizer (nanochat/optim.py, including the Muon optimizer) invite direct modification, enabling rapid prototyping of novel architectures.
Because the entire codebase remains under 2,000 lines of core logic, modifications compile and test in seconds—dramatically accelerating the research iteration cycle compared to heavier frameworks.
Why nanochat Represents the Future of Accessible LLM Development
In an era of increasingly complex LLM tooling, nanochat returns focus to fundamentals: minimal code, maximal insight, and compute efficiency. Its single-GPU compatibility and sub-$100 training cost remove traditional barriers, allowing individual developers and small teams to conduct frontier experiments that once required institutional resources.
By delivering a complete, production-ready chat pipeline in a hackable format, nanochat bridges the gap between educational nanoGPT exercises and real-world deployment. The framework’s integration of pre-training, fine-tuning, and interactive UI creates a self-contained ecosystem for building custom LLMs tailored to specific domains.
For any practitioner seeking to master LLM training without complexity, nanochat provides the clearest path forward. Clone the repository, run the speedrun, and within hours you will possess a fully functional, open-source ChatGPT equivalent trained on your own hardware—proving that powerful LLM development no longer requires massive clusters or endless configuration.












