DeerFlow: The Open-Source Super Agent Harness That Researches, Codes, and Creates
Key Takeaways: DeerFlow 2.0 is a free, MIT-licensed super agent harness from ByteDance that orchestrates parallel sub-agents inside isolated Docker sandboxes with persistent memory, extensible skills, and multi-model LLM support — giving developers a batteries-included runtime for tasks that take minutes to hours.
What Is DeerFlow?
DeerFlow — short for Deep Exploration and Efficient Research Flow — is an open-source super agent harness developed by ByteDance and released under the MIT license. On February 28, 2026, DeerFlow 2.0 claimed the number-one spot on GitHub Trending following a complete ground-up rewrite of the project. The official site is deerflow.tech.
DeerFlow is not a chatbot wrapper or a thin prompt orchestration layer. It is a full-stack agent runtime built on LangGraph and LangChain that provides everything an agent needs to get real work done: a file system, long-term memory, a sandboxed execution environment, modular skills, extensible tools, and the ability to spawn and coordinate parallel sub-agents.
DeerFlow 2.0 is a ground-up rewrite and shares no code with version 1. If you need the original deep research framework, it is preserved on the 1.x branch. Active development has fully moved to 2.0.

From Deep Research Framework to Super Agent Harness
DeerFlow started as a focused deep research tool. The community quickly pushed it into territories the original authors had not anticipated: building data pipelines, generating slide decks, spinning up dashboards, automating content workflows. That breadth revealed that DeerFlow was not just a research tool — it was a harness, a runtime that gives agents the infrastructure to get work done.
The 2.0 rewrite formalizes that insight. DeerFlow now ships batteries included: a filesystem, persistent memory, sandboxed execution, pre-built skills, and multi-agent orchestration ready to deploy with a single make docker-start. The project remains fully extensible — every built-in skill, tool, and model can be swapped out or supplemented.
Core Features
Skills and Tools
Skills are structured capability modules defined as Markdown files. Each skill describes a workflow, best practices, and references to supporting resources. DeerFlow ships with built-in skills for research, report generation, slide creation, web page building, and image generation. You can add custom skills in /mnt/skills/custom/ without touching the core codebase.
Skills are loaded progressively — only when the agent determines a task needs them. This keeps context windows lean and makes DeerFlow viable on token-sensitive or smaller models.
The built-in toolset covers web_search (Tavily), web_fetch (Jina AI), ls, read_file, write_file, str_replace, and bash execution. Additional tools can be added via MCP servers or plain Python functions.
Sub-Agents
Complex tasks rarely fit in a single agent pass. The lead agent decomposes tasks and spawns sub-agents on the fly, each with its own scoped context, tool access, and termination conditions. Sub-agents run in parallel when the task permits and return structured results that the lead agent synthesizes into a coherent final output.
A research task might fan out into a dozen sub-agents, each exploring a different angle, then converge into a single report, website, or slide deck with generated visuals — all orchestrated by a single harness.
Sandbox and File System
Each task runs inside an isolated Docker container with a full filesystem. The agent reads, writes, and edits files; executes bash commands; and runs code. All activity is sandboxed and auditable, with zero contamination between sessions.
/mnt/user-data/
uploads/ # Your input files
workspace/ # Agent working directory
outputs/ # Final deliverablesThe All-in-One Sandbox combines browser, shell, file system, MCP, and VSCode Server in a single Docker container and is recommended for production deployments.
Context Engineering
Each sub-agent runs in its own isolated context, unable to see the state of the main agent or sibling sub-agents. Within a session, DeerFlow manages the overall context window aggressively: completed sub-tasks are summarized, intermediate results are offloaded to the filesystem, and no-longer-relevant content is compressed. This keeps the harness sharp across long, multi-step tasks without blowing token budgets.
Long-Term Memory
DeerFlow builds a persistent memory of user preferences, writing style, technical stack, and recurring workflows across sessions. Memory is stored locally and remains entirely under user control. The more the system is used, the more context it carries forward — improving output quality over time without requiring external vector databases or cloud sync.
Installation Guide
Prerequisites
For local development, the following must be installed and available on the system PATH:
- Node.js 22+ and pnpm
- uv (Python package manager)
- nginx (used as a local reverse proxy)
- Docker (required for sandbox isolation; also required for the Docker deployment path)
Run make check in the project root to verify all prerequisites are met.
Option 1: Docker (Recommended)
Docker is the fastest path to a consistent, production-grade environment.
# Clone the repository
git clone https://github.com/bytedance/deer-flow.git
cd deer-flow
# Generate local configuration files from templates
make config
# Edit config.yaml to define at least one model and set API keys
# Edit .env to add TAVILY_API_KEY, OPENAI_API_KEY, etc.
# Pull the sandbox image (run once, or when the image updates)
make docker-init
# Start all services
make docker-startAccess the interface at http://localhost:2026.
Option 2: Local Development
git clone https://github.com/bytedance/deer-flow.git
cd deer-flow
# Verify prerequisites
make check
# Install backend and frontend dependencies
make install
# (Optional) Pre-pull the Docker sandbox image
make setup-sandbox
# Start all services
make devAccess the interface at http://localhost:2026.
Configuration
Model Configuration
Edit config.yaml in the project root to define one or more LLM providers. DeerFlow supports any LangChain-compatible provider, including OpenAI, Anthropic, DeepSeek, and OpenAI-compatible gateways.
models:
- name: gpt-4
display_name: GPT-4
use: langchain_openai:ChatOpenAI
model: gpt-4
api_key: $OPENAI_API_KEY
max_tokens: 4096
temperature: 0.7API keys should always be referenced via environment variables (the $ prefix). Set them in the .env file at the project root:
OPENAI_API_KEY=your-key-here
TAVILY_API_KEY=your-key-hereDeerFlow performs best with models that support long context windows (100k+ tokens), strong reasoning, multimodal input, and robust tool-use. Supported providers include OpenAI, Anthropic, DeepSeek, Doubao, Gemini, and any OpenAI-compatible API gateway.
Sandbox Configuration
Three sandbox modes are available:
# Local (simpler setup, runs code on host)
sandbox:
use: src.sandbox.local:LocalSandboxProvider
# Docker (isolated containers, recommended for production)
sandbox:
use: src.community.aio_sandbox:AioSandboxProvider
# Docker + Kubernetes (isolated pods via provisioner service)
sandbox:
use: src.community.aio_sandbox:AioSandboxProvider
provisioner_url: http://provisioner:8002For production, the Docker sandbox provides the best balance of isolation and practicality. Kubernetes mode is available for teams running existing cluster infrastructure.
Skills Configuration
Skills are stored in deer-flow/skills/{public,custom}/. Each skill requires a SKILL.md file with metadata. Custom skills placed in the custom/ directory are automatically discovered and made available to the agent without any additional registration steps.
Using DeerFlow
After starting the interface at http://localhost:2026, open the chat panel and submit a task in plain language. DeerFlow plans the approach, spawns the necessary sub-agents, executes code or web searches inside the sandbox, and returns structured output — files, reports, slides, or code — directly to the interface.
The embedded Python client supports programmatic access without running the full HTTP stack:
from src.client import DeerFlowClient
client = DeerFlowClient()
response = client.chat("Analyze this dataset and generate a summary report", thread_id="task-001")
for event in client.stream("hello"):
if event.type == "messages-tuple" and event.data.get("type") == "ai":
print(event.data["content"])Use Cases
DeerFlow’s architecture supports a wide range of autonomous task categories:
- Academic and market research: Multi-source deep research synthesis across dozens of documents and live web data, delivered as a structured report.
- Software development assistance: Codebase analysis, automated refactoring, test generation, and documentation across large file trees.
- Content and media production: Slide deck generation, web page creation, image generation, and long-form report writing — all within a single agent session.
- Data pipeline automation: Continuous data ingestion, transformation, and reporting pipelines that persist state across multi-hour runs.
- Competitive intelligence: Scheduled agent runs that monitor topics, aggregate findings, and surface signals across structured and unstructured sources.
- Internal knowledge tooling: Organizations can deploy DeerFlow as a self-hosted knowledge worker that retains institutional context across sessions using long-term memory.
Development and Extension Ideas
DeerFlow’s architecture is built for extension. The following directions are immediately tractable for teams looking to go beyond the defaults:
- Custom skill libraries: Write
SKILL.mdfiles for domain-specific workflows — legal document review, financial modeling, medical literature synthesis — and mount them into the skills directory. - MCP server integration: Connect DeerFlow to proprietary APIs, internal databases, or enterprise tools via the MCP server protocol without modifying core agent logic.
- Multi-tenant SaaS deployment: Combine DeerFlow’s per-session sandbox isolation with a reverse proxy and authentication layer to build a shared deployment with isolated workspaces per user.
- Automated reporting pipelines: Wire DeerFlow’s Python client into scheduled jobs (cron, Temporal, Airflow) to generate periodic reports, digests, or briefings.
- SIEM and threat intelligence feeds: Use DeerFlow’s web fetch and bash tools to monitor security data sources, correlate events, and push structured findings to downstream systems.
- Slide and dashboard generation workflows: Extend the built-in slide-creation skill to pull live data from APIs, build branded templates, and publish outputs to cloud storage.
Conclusion
DeerFlow 2.0 is one of the most capable open-source agent runtimes available today. Its combination of isolated Docker sandboxes, parallel sub-agent orchestration, progressive skill loading, and persistent long-term memory addresses the real engineering challenges that autonomous agents face at production scale. The MIT license, the self-hosted deployment model, and the model-agnostic design make it a practical choice for developers, researchers, and organizations that need full control over their agent infrastructure.
The project is active, growing rapidly, and open to contributions. Explore the codebase, run the quickstart, and try deploying a custom skill to understand what DeerFlow 2.0 makes possible.
- Repository: https://github.com/bytedance/deer-flow
- Official Website: https://deerflow.tech/
- Configuration Guide: https://github.com/bytedance/deer-flow/blob/main/backend/docs/CONFIGURATION.md
- Contributing Guide: https://github.com/bytedance/deer-flow/blob/main/CONTRIBUTING.md
- License: MIT








