A Pi-based coding agent with local multi-session and real-time voice interfaces.
New issues and PRs from new contributors are auto-closed by default. Maintainers review auto-closed issues daily. See CONTRIBUTING.md.
The voice agent turns the coding harness into a browser-based, embodied assistant. The demo above shows the actual local WebUI and a live Hanhan session. Click it to open the MP4 recording.
- Real-time voice loop: architecture-specific VAD, Tencent ASR, text/audio semantic endpoint detection, and MiniMax streaming TTS run as a TEN Framework graph.
- Animated 3D presence: Three.js renders a bundled character with idle motion, blinking, pointer tracking, gestures, and speech-driven mouth movement.
- Coding-agent access: recognized speech and text input enter a normal Hanhan
AgentSession, including its workspace and tools. - Local browser transport: microphone capture, transcript rendering, and PCM playback stay in a local WebUI without Agora RTC.
- Natural interruption: speech can abort an active response and invalidate queued LLM, TTS, subtitle, and browser playback output before the next turn.
The browser connects to a standalone manager that owns configuration and session lifecycle. Each session starts a TEN backend containing the real-time graph, the main_nodejs control plane, and an in-process Hanhan AgentSession.
Uplink PCM branches into VAD/ASR and optional audio endpoint inference. Accepted turns enter Hanhan once, then stream through sentence buffering, MiniMax TTS, browser playback, and timestamp-aligned captions.
main_nodejs serializes input events, rejects stale endpoint results, applies the configured barge-in policy, and fences every output by turn ID. Interruption invalidates the old turn before aborting the agent and flushing TTS and browser playback.
| Extension | Runtime | Responsibility |
|---|---|---|
webui |
Node.js | Bridges browser WebSocket PCM/JSON with TEN audio, data, commands, transcripts, state, and metrics. |
streamid_adapter |
TEN | Normalizes the microphone stream ID before audio processing. |
speaker_gate |
Python | Optionally applies local CAM++ speaker-conditioned gating before VAD, ASR, and Smart Turn. |
vad |
Python | Emits speech start/end events; x64 uses TEN VAD and arm64 uses Silero VAD. |
stt |
Python | Streams Tencent ASR partial/final results; final text is the voice prompt source. |
semantic_eos |
Python | Judges turn completeness from finalized text and bounded conversation history. |
smart_turn |
Python | Judges turn completeness directly from gated PCM using Smart Turn v3.2. |
main_control |
Node.js | Owns turns, endpoint acceptance, barge-in, Hanhan prompts, output ordering, flushes, and metrics. |
tts |
Python | Streams MiniMax PCM plus word timestamps for playback-aligned captions. |
See the voice-agent example for setup, architecture, service requirements, and asset attribution.
Hanhan is a focused fork of Pi, not a separate agent runtime. It keeps Pi's provider abstraction, tool calling, session model, extension system, and terminal UI, then adds a product-facing layer around them.
| Area | Upstream Pi | Hanhan |
|---|---|---|
| CLI | pi |
Renamed and packaged as hanhan |
| Primary interface | Terminal coding agent | Terminal plus a local multi-session WebUI |
| Voice | No integrated real-time voice example | TEN-based ASR/VAD/TTS pipeline with barge-in |
| Visual presence | Terminal UI | Three.js 3D avatar synchronized with agent state and audio |
| Distribution | Upstream Pi release flow | Hanhan-branded Node and Bun release artifacts |
Changes that are useful to the wider Pi ecosystem should remain compatible with the upstream architecture where practical. Upstream Pi documentation is still relevant for the shared core.
Hanhan builds on the Pi agent harness packages, including its self-extensible coding agent:
- @earendil-works/pi-coding-agent: Interactive coding agent CLI
- @earendil-works/pi-agent-core: Agent runtime with tool calling and state management
- @earendil-works/pi-ai: Unified multi-provider LLM API (OpenAI, Anthropic, Google, …)
To learn more about the shared Pi foundation:
- Visit pi.dev, the project website with demos
- Read the documentation, but you can also ask the agent to explain itself
| Package | Description |
|---|---|
| @earendil-works/pi-ai | Unified multi-provider LLM API (OpenAI, Anthropic, Google, etc.) |
| @earendil-works/pi-agent-core | Agent runtime with tool calling and state management |
| @earendil-works/pi-coding-agent | Interactive coding agent CLI |
| @earendil-works/pi-tui | Terminal UI library with differential rendering |
For Slack/chat automation and workflows see earendil-works/pi-chat.
Pi does not include a built-in permission system for restricting filesystem, process, network, or credential access. By default, it runs with the permissions of the user and process that launched it.
If you need stronger boundaries, containerize or sandbox Pi. See packages/coding-agent/docs/containerization.md for three patterns:
- Gondolin extension: keep
piand provider auth on the host while routing built-in tools and!commands into a local Linux micro-VM. - Plain Docker: run the whole
piprocess in a local container for simple isolation. - OpenShell: run the whole
piprocess in a policy-controlled sandbox.
See CONTRIBUTING.md for contribution guidelines and AGENTS.md for project-specific rules (for both humans and agents). Longer term plans for Pi can also be found in RFCs.
npm install --ignore-scripts # Install all dependencies without running lifecycle scripts
npm run build # Refresh model data, then build all packages
npm run build:offline # Rebuild using existing model data without network access
npm run check # Lint, format, and type check
./test.sh # Run tests (skips LLM-dependent tests without API keys)
./pi-test.sh # Run pi from sources (can be run from any directory)GitHub releases include a versioned source archive covered by the release's SHA256SUMS file. Extract it and run the same build script used for the official standalone binaries:
VERSION="<release-version>"
tar -xzf "pi-${VERSION}-source.tar.gz"
cd "pi-${VERSION}"
./scripts/build-binaries.sh --offline-model-data --platform linux-x64 --out "$PWD/out"The source archive includes the generated provider model data used for the release. --offline-model-data builds with that snapshot instead of refreshing it from live provider catalogs. The script still installs dependencies, builds the monorepo, compiles the Bun executable, and stages its runtime assets. Package maintainers who provide dependencies separately can pass --skip-install --skip-deps.
We treat npm dependency changes as reviewed code changes.
- Direct external dependencies are pinned to exact versions. Internal workspace packages remain version-ranged.
.npmrcsetssave-exact=trueandmin-release-age=2to avoid same-day dependency releases during npm resolution.package-lock.jsonis the dependency ground truth. Pre-commit blocks accidental lockfile commits unlessPI_ALLOW_LOCKFILE_CHANGE=1is set.npm run checkverifies pinned direct deps, native TypeScript import compatibility, and the generated coding-agent shrinkwrap.- The published CLI package includes
packages/coding-agent/npm-shrinkwrap.json, generated from the root lockfile, to pin transitive deps for npm users. - Release smoke tests use
npm run release:localto build, pack, and create isolated npm and Bun installs outside the repo before tagging a release. - Local release installs, documented npm installs, and
pi update --selfuse--ignore-scriptswhere supported. - CI installs with
npm ci --ignore-scripts, and a scheduled GitHub workflow runsnpm audit --omit=devplusnpm audit signatures --omit=dev. - Shrinkwrap generation has an explicit allowlist for dependency lifecycle scripts; new lifecycle-script deps fail checks until reviewed.
If you use Pi or other coding agents for open source work, please share your sessions.
Public OSS session data helps improve coding agents with real-world tasks, tool use, failures, and fixes instead of toy benchmarks.
For the full explanation, see this post on X.
To publish sessions, use badlogic/pi-share-hf. Read its README.md for setup instructions. All you need is a Hugging Face account, the Hugging Face CLI, and pi-share-hf.
You can also watch this video, where I show how I publish my pi-mono sessions.
I regularly publish my own pi-mono work sessions here:
MIT



