AI & AUTOMATION

Ouroboros: The Open-Source Agent OS That Turns Vague Prompts Into Verified, Replayable AI Coding Workflows

Key Takeaways: Ouroboros is an open-source, local-first Agent OS that replaces ad-hoc AI prompting with a Socratic, specification-first loop (interview, seed, execute, evaluate, evolve), so you get replayable, policy-bound, multi-runtime AI coding workflows that converge on a verified codebase instead of an endless prompt-and-pray cycle.

If you have ever shipped a feature with Claude Code or Codex CLI only to discover, three commits later, that the agent quietly invented an assumption you never made, you already understand the problem Ouroboros is trying to solve. Most AI coding tools fail at the input, not the output. They are extraordinary code generators wired to under-specified prompts, and that gap is where rework, drift, and silent regressions live.

Ouroboros, the project hosted at github.com/Q00/ouroboros, reframes the problem. Instead of treating the LLM as a faster typist, it treats agent work as something that needs an operating system: a stable kernel of primitives that enforces specifications, records every event, and runs the same workflow across multiple AI coding CLIs.

What Is Ouroboros

Ouroboros calls itself an Agent OS: a local-first runtime layer that turns non-deterministic agent work into a replayable, observable, policy-bound execution contract. Its tagline, “Stop prompting. Start specifying,” is the whole thesis in four words.

The project ships as three coordinated repositories that form a familiar stack:

LayerRepoRole
Shell (terminal client)Q00/ourocodeNative TUI cockpit that surfaces Ouroboros state across CLIs
Apps (domain workflows)Q00/ouroboros-pluginsInstallable plugins for PR ops, Jira sync, incidents, releases
OS (this repo)Q00/ouroborosThe ooo commands, spec-first workflow engine, multi-runtime adapter

The kernel is what most users interact with through the ooo command set. Plugins compose its primitives into scoped, auditable domain programs, and ourocode is the optional unified terminal interface.

Core Concepts In Plain English

  • Seed. An immutable specification produced from a Socratic interview. Nothing is built until the seed exists, and the seed is what locks intent.
  • Ledger. Every action becomes a Seed-bound, ledger-recorded, replayable event, regardless of which LLM executed it. This is what makes runs reproducible.
  • Runtime adapter. A pluggable backend layer that lets the same workflow execute on Claude Code, Codex CLI, OpenCode, Hermes, Gemini CLI, Kiro CLI, or GitHub Copilot CLI.
  • The Loop. Interview, Seed, Execute, Evaluate, Evolve, then feed evaluation back into the next generation’s seed. The serpent devours its own tail; each cycle the system knows more than the last.
  • Ambiguity Score. A weighted, LLM-scored metric (Goal, Constraint, Success Criteria, and more). Code generation is gated until ambiguity drops to roughly 0.2 or below.

Why It Matters

Problem with vanilla AI codingWhat Ouroboros does
Vague prompts cause the agent to guessSocratic interview surfaces hidden assumptions before any code is written
No spec, so architecture drifts mid-buildImmutable Seed locks intent and acceptance criteria
“Looks good” is not verificationA three-stage gate runs Mechanical (free), then Semantic, then Multi-Model Consensus
Sessions die when your laptop sleepsEvent sourcing reconstructs the full lineage so ooo ralph resumes after restarts

How to Install Ouroboros

The project requires Python 3.12 or newer. Beyond that, the installer is genuinely one line.

Option 1: One-Line Installer (Recommended)

curl -fsSL https://raw.githubusercontent.com/Q00/ouroboros/main/scripts/install.sh | bash

The script auto-detects Claude Code, Codex CLI, and Hermes CLI on your machine, installs the ouroboros-ai Python package, and registers the Ouroboros MCP server with whichever runtimes are present.

Option 2: pip, uv, or pipx

If you prefer to control the install yourself:

# Base package
pip install ouroboros-ai

# With extras (pick what you need)
pip install 'ouroboros-ai[claude]'    # Claude Code integration
pip install 'ouroboros-ai[litellm]'   # 100+ providers via LiteLLM
pip install 'ouroboros-ai[mcp]'       # MCP server and client
pip install 'ouroboros-ai[tui]'       # Textual terminal UI
pip install 'ouroboros-ai[all]'       # Everything

After installation, configure your runtime:

ouroboros setup

For runtimes the installer cannot auto-detect, pass an explicit flag:

ouroboros setup --runtime opencode    # or kiro, copilot

Option 3: Claude Code Plugin Only

If you do not want a system-level Python install, you can run Ouroboros entirely as a Claude Code plugin:

claude plugin marketplace add Q00/ouroboros && \
claude plugin install ouroboros@ouroboros

Then run ooo setup inside a Claude Code session.

Verifying the Install

Open the AI coding agent you registered (Claude Code, Codex CLI, Copilot CLI, etc.) and run:

> ooo help

You should see the full skill catalogue: interview, seed, run, evaluate, evolve, ralph, unstuck, status, and friends. From a plain terminal, the equivalent CLI is ouroboros --help.

To uninstall cleanly later, ouroboros uninstall removes the package, configuration, MCP registration, and local data.

Using Ouroboros: From Vague Idea to Verified Codebase

The intended flow is simple to describe and surprisingly disciplined in practice.

Step 1: Interview

Start with a one-sentence idea, no matter how rough:

> ooo interview "I want to build a task management CLI"

Ouroboros launches the Socratic Interviewer agent. It refuses to write code and only asks questions, scoring the ambiguity of your answers as it goes. By the time the interview finishes, hidden assumptions are explicit and the ambiguity score is below the gating threshold.

From the terminal, the same step is:

ouroboros init start

Step 2: Seed

When the interview converges, your answers crystallize into an immutable seed specification (a YAML file with acceptance criteria, ontology, and constraints). This is the contract every later step is bound to.

Step 3: Run

Execute the seed:

ouroboros run seed.yaml

Internally, Ouroboros decomposes the work using a Double Diamond (Discover, Define, Design, Deliver) and routes each task through the PAL Router, a three-tier cost optimizer that escalates from frugal models to frontier models only when needed and downgrades again on success.

Step 4: Evaluate

The 3-stage gate runs automatically: a free Mechanical check (lint, build, tests), then Semantic evaluation, then Multi-Model Consensus. “Looks good” is replaced by a verifiable verdict.

Step 5: Evolve and Ralph

If the evaluator finds gaps, ooo evolve feeds the result back as input to the next generation’s seed. For long-running, stateless persistence, ooo ralph keeps the loop alive across sessions and restarts, stopping only when the ontology stabilizes (similarity >= 0.95).

Useful Day-to-Day Commands

SkillWhat it does
ooo unstuckLoads five lateral-thinking personas when you are blocked
ooo statusLists in-flight sessions and detects drift
ooo resume-sessionReattaches to an interrupted workflow
ooo cancelCancels stuck or orphaned executions
ooo pmRuns a PM-focused interview and generates a PRD
ooo brownfieldScans an existing repository and configures sensible defaults
ooo publishPublishes a Seed as GitHub Epic and Task issues for team workflows

A Note on Brownfield Projects

Ouroboros is not just for greenfield ideas. The bigbang.brownfield_explorer auto-detects config files across multiple language ecosystems, and the ambiguity weights shift (35% Goal, 25% Constraint, and so on) to reflect the reality that an existing codebase already encodes many decisions.


Ouroboros is a deliberately opinionated answer to a problem the AI coding community has mostly tried to solve with better prompts: agents do not fail because they cannot code, they fail because they were never told precisely what to build. By installing a Seed, a Ledger, and an Ambiguity Gate between you and the model, the project turns “vibe coding” into something closer to engineering, while staying agnostic about which CLI or model you prefer.

If you already live inside Claude Code, Codex CLI, or any of the supported runtimes, the cost of trying it is one curl command and a fifteen-minute interview. The reward is a workflow you can replay, audit, and evolve, instead of one you have to re-prompt every Monday. Spin it up on a side project this week, watch the ambiguity score drop in real time, and see how much rework quietly disappears.

You may also like

Subscribe
Notify of
guest

0 Comments
Newest
Oldest Most Voted