Key Takeaways: Ouroboros is an open-source, local-first Agent OS that replaces ad-hoc AI prompting with a Socratic, specification-first loop (interview, seed, execute, evaluate, evolve), so you get replayable, policy-bound, multi-runtime AI coding workflows that converge on a verified codebase instead of an endless prompt-and-pray cycle.
If you have ever shipped a feature with Claude Code or Codex CLI only to discover, three commits later, that the agent quietly invented an assumption you never made, you already understand the problem Ouroboros is trying to solve. Most AI coding tools fail at the input, not the output. They are extraordinary code generators wired to under-specified prompts, and that gap is where rework, drift, and silent regressions live.
Ouroboros, the project hosted at github.com/Q00/ouroboros, reframes the problem. Instead of treating the LLM as a faster typist, it treats agent work as something that needs an operating system: a stable kernel of primitives that enforces specifications, records every event, and runs the same workflow across multiple AI coding CLIs.
What Is Ouroboros
Ouroboros calls itself an Agent OS: a local-first runtime layer that turns non-deterministic agent work into a replayable, observable, policy-bound execution contract. Its tagline, “Stop prompting. Start specifying,” is the whole thesis in four words.
The project ships as three coordinated repositories that form a familiar stack:
| Layer | Repo | Role |
|---|---|---|
| Shell (terminal client) | Q00/ourocode | Native TUI cockpit that surfaces Ouroboros state across CLIs |
| Apps (domain workflows) | Q00/ouroboros-plugins | Installable plugins for PR ops, Jira sync, incidents, releases |
| OS (this repo) | Q00/ouroboros | The ooo commands, spec-first workflow engine, multi-runtime adapter |
The kernel is what most users interact with through the ooo command set. Plugins compose its primitives into scoped, auditable domain programs, and ourocode is the optional unified terminal interface.

Core Concepts In Plain English
- Seed. An immutable specification produced from a Socratic interview. Nothing is built until the seed exists, and the seed is what locks intent.
- Ledger. Every action becomes a Seed-bound, ledger-recorded, replayable event, regardless of which LLM executed it. This is what makes runs reproducible.
- Runtime adapter. A pluggable backend layer that lets the same workflow execute on Claude Code, Codex CLI, OpenCode, Hermes, Gemini CLI, Kiro CLI, or GitHub Copilot CLI.
- The Loop. Interview, Seed, Execute, Evaluate, Evolve, then feed evaluation back into the next generation’s seed. The serpent devours its own tail; each cycle the system knows more than the last.
- Ambiguity Score. A weighted, LLM-scored metric (Goal, Constraint, Success Criteria, and more). Code generation is gated until ambiguity drops to roughly 0.2 or below.
Why It Matters
| Problem with vanilla AI coding | What Ouroboros does |
|---|---|
| Vague prompts cause the agent to guess | Socratic interview surfaces hidden assumptions before any code is written |
| No spec, so architecture drifts mid-build | Immutable Seed locks intent and acceptance criteria |
| “Looks good” is not verification | A three-stage gate runs Mechanical (free), then Semantic, then Multi-Model Consensus |
| Sessions die when your laptop sleeps | Event sourcing reconstructs the full lineage so ooo ralph resumes after restarts |
How to Install Ouroboros
The project requires Python 3.12 or newer. Beyond that, the installer is genuinely one line.
Option 1: One-Line Installer (Recommended)
curl -fsSL https://raw.githubusercontent.com/Q00/ouroboros/main/scripts/install.sh | bashThe script auto-detects Claude Code, Codex CLI, and Hermes CLI on your machine, installs the ouroboros-ai Python package, and registers the Ouroboros MCP server with whichever runtimes are present.
Option 2: pip, uv, or pipx
If you prefer to control the install yourself:
# Base package
pip install ouroboros-ai
# With extras (pick what you need)
pip install 'ouroboros-ai[claude]' # Claude Code integration
pip install 'ouroboros-ai[litellm]' # 100+ providers via LiteLLM
pip install 'ouroboros-ai[mcp]' # MCP server and client
pip install 'ouroboros-ai[tui]' # Textual terminal UI
pip install 'ouroboros-ai[all]' # EverythingAfter installation, configure your runtime:
ouroboros setupFor runtimes the installer cannot auto-detect, pass an explicit flag:
ouroboros setup --runtime opencode # or kiro, copilotOption 3: Claude Code Plugin Only
If you do not want a system-level Python install, you can run Ouroboros entirely as a Claude Code plugin:
claude plugin marketplace add Q00/ouroboros && \
claude plugin install ouroboros@ouroborosThen run ooo setup inside a Claude Code session.
Verifying the Install
Open the AI coding agent you registered (Claude Code, Codex CLI, Copilot CLI, etc.) and run:
> ooo helpYou should see the full skill catalogue: interview, seed, run, evaluate, evolve, ralph, unstuck, status, and friends. From a plain terminal, the equivalent CLI is ouroboros --help.
To uninstall cleanly later, ouroboros uninstall removes the package, configuration, MCP registration, and local data.
Using Ouroboros: From Vague Idea to Verified Codebase
The intended flow is simple to describe and surprisingly disciplined in practice.
Step 1: Interview
Start with a one-sentence idea, no matter how rough:
> ooo interview "I want to build a task management CLI"Ouroboros launches the Socratic Interviewer agent. It refuses to write code and only asks questions, scoring the ambiguity of your answers as it goes. By the time the interview finishes, hidden assumptions are explicit and the ambiguity score is below the gating threshold.
From the terminal, the same step is:
ouroboros init startStep 2: Seed
When the interview converges, your answers crystallize into an immutable seed specification (a YAML file with acceptance criteria, ontology, and constraints). This is the contract every later step is bound to.
Step 3: Run
Execute the seed:
ouroboros run seed.yamlInternally, Ouroboros decomposes the work using a Double Diamond (Discover, Define, Design, Deliver) and routes each task through the PAL Router, a three-tier cost optimizer that escalates from frugal models to frontier models only when needed and downgrades again on success.
Step 4: Evaluate
The 3-stage gate runs automatically: a free Mechanical check (lint, build, tests), then Semantic evaluation, then Multi-Model Consensus. “Looks good” is replaced by a verifiable verdict.
Step 5: Evolve and Ralph
If the evaluator finds gaps, ooo evolve feeds the result back as input to the next generation’s seed. For long-running, stateless persistence, ooo ralph keeps the loop alive across sessions and restarts, stopping only when the ontology stabilizes (similarity >= 0.95).
Useful Day-to-Day Commands
| Skill | What it does |
|---|---|
ooo unstuck | Loads five lateral-thinking personas when you are blocked |
ooo status | Lists in-flight sessions and detects drift |
ooo resume-session | Reattaches to an interrupted workflow |
ooo cancel | Cancels stuck or orphaned executions |
ooo pm | Runs a PM-focused interview and generates a PRD |
ooo brownfield | Scans an existing repository and configures sensible defaults |
ooo publish | Publishes a Seed as GitHub Epic and Task issues for team workflows |
A Note on Brownfield Projects
Ouroboros is not just for greenfield ideas. The bigbang.brownfield_explorer auto-detects config files across multiple language ecosystems, and the ambiguity weights shift (35% Goal, 25% Constraint, and so on) to reflect the reality that an existing codebase already encodes many decisions.
Ouroboros is a deliberately opinionated answer to a problem the AI coding community has mostly tried to solve with better prompts: agents do not fail because they cannot code, they fail because they were never told precisely what to build. By installing a Seed, a Ledger, and an Ambiguity Gate between you and the model, the project turns “vibe coding” into something closer to engineering, while staying agnostic about which CLI or model you prefer.
If you already live inside Claude Code, Codex CLI, or any of the supported runtimes, the cost of trying it is one curl command and a fifteen-minute interview. The reward is a workflow you can replay, audit, and evolve, instead of one you have to re-prompt every Monday. Spin it up on a side project this week, watch the ambiguity score drop in real time, and see how much rework quietly disappears.







