All agent CLIs
PenguinHarness in NeuroSquad

PenguinHarness,with a canvas around it

Run as many Penguin agents as you need, each the real penguin chat in its own card. NeuroSquad runs its own Penguin server and gives every card an agent of its own; Penguin’s status stream tells the card when it works, waits or finishes, and arrows hand it terminals, notes and other agents.

Official sitepenguin.ooo
  • The real penguin chat, your own models
  • Your ~/.penguin is never written
  • Free, for Windows and macOS
Status stream and hook

It knows when Penguin is done — and when it’s asking

PenguinHarness is a client and a server. NeuroSquad listens to the server’s own status stream for every card’s session, and a small hook package in each card’s agent decides which tool calls need you.

session_state: runningsignal
Working

The server’s stream reports the card’s session running. The card’s icon spins, the Squad status card counts it, and the toast from its last question is taken down.

pre_tool_usesignal
Needs you

The card’s pre_tool_use hook leaves a call to Penguin’s own “? Approve this tool call? [Y/n]”. The toast names it — PenguinHarness needs your permission: exec_command: rm -rf dist — until you answer in the card.

session_state: idlesignal
Finished

The stream reports idle and the chat is back at its > prompt. The next queued prompt goes out, the journal gets the answer, and a toast says it’s done.

The hook is the policy

The card runs Penguin with --approve always-ask, and its hook answers first: NeuroSquad’s own tools, read_file and subagents run unasked, as in Claude Code; everything else waits on you.

Subagents show up too

A subagent runs in a child session; the stream reports it, so the card stays working while it runs and Usage lists it apart.

No early “finished”

“Finished” waits until the chat is back at its prompt — otherwise a queued prompt pasted a moment too early would be lost.

Watch it work

Three things you’ll do with PenguinHarness

Replicas of the app, played step by step. When Penguin asks, the answer is yours.

Penguin runs the tests through a terminal card and reads the file unasked, stops on its own approval question before editing, and finishes — the journal writes itself.

Everything it gets

Everything PenguinHarness can do in NeuroSquad

Measured on PenguinHarness 0.2.13 and on its main branch. Some of it is Penguin’s own — MCP, sessions, its Trace files — and much is built around its server. Where it can’t be done, the card says so.

  • 01

    Knows its state

    Penguin’s own status stream and the card’s hook report every turn and every call that needs you.

  • Built by NeuroSquad

    Finished and waiting are different

    PenguinHarness 0.2.13 has no per-prompt hook, so the card follows the server’s own status stream (session_state running and idle) for its session — every turn, including ones started from the phone.

    Docs: Finished and waiting are different
  • Built by NeuroSquad

    A notification that says what it needs

    The toast names the call — PenguinHarness needs your permission: write_file: a.txt — and the card goes back to working when you answer it in the card.

    Docs: A notification that says what it needs
  • Built by NeuroSquad

    Stops through its server

    Stop, the budget brake and the phone’s Stop call Penguin’s own abort API. Ctrl-C on an idle chat would ask to quit, so Stop never sends it, and the phone’s key row marks Ctrl+C as quitting.

    Docs: Stops through its server
  • Picks up where it left off

    The stream and the hooks report the session id; the next launch is penguin chat --resume <id>. A deleted session starts fresh by itself.

    Docs: Picks up where it left off
  • The Squad status card

    Every AI agent of the workspace on one card, with its status, model and turn time.

    Docs: The Squad status card
  • 02

    Works through the canvas

    An arrow from the card is a set of tools, over Penguin’s own MCP client.

  • Arrows are tools

    The card’s agent lists NeuroSquad’s MCP server in its tools.mcpServers. An arrow to a terminal, note, browser or agent gives Penguin those tools.

    Docs: Arrows are tools
  • Built by NeuroSquad

    The arrow is the consent

    The card’s hook lets NeuroSquad’s own tools, read_file and subagents run without a prompt — what Claude Code runs without asking. Everything else waits on you.

    Docs: The arrow is the consent
  • Built by NeuroSquad

    Tools offered up front

    Penguin takes a snapshot of a model context’s tools and ignores list changes. So a card gets every NeuroSquad capability from the start, and the arrows decide at call time.

    Docs: Tools offered up front
  • Built by NeuroSquad

    Canvas mode that holds

    The hook refuses exec_command and input_command — the shell, and Penguin’s only way to the web — before any approval, dangerous mode included.

    Docs: Canvas mode that holds
  • MCP servers from the registry

    Installed servers come as a second server with no pre-approval: Penguin asks before every call to it.

    Docs: MCP servers from the registry
  • 03

    Built for long runs

    Context, Compact, rate limits, a prompt queue and a journal — handled on the card.

  • Context meter

    The latest request’s tokens from the Trace, against the window Penguin records for the session’s model.

    Docs: Context meter
  • Compact

    The card’s Compact button sends Penguin’s own /compact.

    Docs: Compact
  • Built by NeuroSquad

    Rate limits on the card

    When a request ends on a rate or quota error, the card reads it from the Trace and shows the limit.

    Docs: Rate limits on the card
  • Prompt queue

    Prompts to a busy card wait and go out after the turn, one by one; long and multi-line prompts arrive as one message.

    Docs: Prompt queue
  • Journal and hand-off

    The last answer is read from Penguin’s Trace files: into a connected note, a hand-off, agent_wait, Telegram and custom cards.

    Docs: Journal and hand-off
  • 04

    Works in a team

    An agent per card, subagents, live dangerous mode and a switch to another CLI.

  • Built by NeuroSquad

    An agent per card, its own server

    NeuroSquad runs its own Penguin server for its own data root and stops it at quit. Each card is a Penguin agent of its own, with its own hooks, MCP servers and instructions; your projects and models are copied in, never changed.

    Docs: An agent per card, its own server
  • Subagents, counted apart

    A subagent runs in a child session with its own Trace. The card stays working while it runs, and Usage, Run Stats and Agent Pulse mark its requests as a subagent’s.

    Docs: Subagents, counted apart
  • Built by NeuroSquad

    Dangerous mode, live

    No restart: while the card’s switch is on, its hook allows every call; turn it off and the next call asks again. Canvas mode still wins.

    Docs: Dangerous mode, live
  • Model and CLI switch

    A lead agent can change the model with agent_set_model — in place on newer Penguin; on 0.2.13 a new session starts with Penguin’s own pointer to the old one — or move the conversation to Claude Code and back.

    Docs: Model and CLI switch
  • Isolated worktrees

    Give an agent its own git worktree, so parallel agents never edit the same files.

    Docs: Isolated worktrees
  • 05

    Cost, control and access

    What it spends, what it may spend, which model it runs on, and how you reach it.

  • Exact usage

    Read from Penguin’s Trace: one record per request, cache reads apart, subagents marked. A model with no known price shows “no price”, never $0.

    Docs: Exact usage
  • Budget brake

    Past the workspace limit each Penguin card is stopped through the server’s abort API — the chat stays open — and automatic prompts wait.

    Docs: Budget brake
  • Built by NeuroSquad

    OpenRouter models

    A provider group of the card’s own goes through NeuroSquad’s local relay, which adds the key and the attribution; the file holds only a placeholder. Your own OpenRouter group is routed the same way.

    Docs: OpenRouter models
  • Built by NeuroSquad

    Your own servers

    Each of your providers becomes a group behind the same relay — OpenAI Chat Completions or Anthropic Messages — with no key on disk.

    Docs: Your own servers
  • From your phone

    Scan a QR code and the whole canvas opens in the phone’s browser: read Penguin’s chat, answer y or n, send the next prompt.

    Docs: From your phone
  • 06

    Good to know

    What Penguin does differently, and what can’t be done.

  • Not possible

    The card can’t pick the session id

    Penguin always creates its own ids. NeuroSquad catches the id from the stream and resumes with it, which works the same in practice.

    Docs: The card can’t pick the session id
  • Not possible

    No cache-write or reasoning split

    Penguin’s Trace folds cache creation into the input and reasoning into the output, so those columns read “not reported”. On Claude models the price is a little low for it.

    Docs: No cache-write or reasoning split
  • Not possible

    No token saver

    Penguin’s pre_tool_use hook can only allow or deny a call, not rewrite it — so RTK can’t work there.

    Docs: No token saver
  • Plugin doors need a newer Penguin

    Caveman, Memory, Context7 and Graphify add text through a per-prompt hook that Penguin runs on every prompt only after 0.2.13. On 0.2.13 their tools still work by arrow.

    Docs: Plugin doors need a newer Penguin
  • This computer only

    The Penguin server would have to run inside a WSL or SSH target, so Penguin cards run on this computer only for now.

    Docs: This computer only
At a glance

Who’s waiting, and what it used

Two cards that answer the questions you ask most when several Penguin agents run at once.

The Squad status card

Every agent of the workspace with its status, model and turn time. The ones waiting on you come first, with the call they want to make. Tap a status to filter.

  • coupon-fixPenguinHarnessdeepseek/deepseek-chatPenguinHarness needs your permission: exec_command: rm -rf distNeeds you2:34
  • api-migrationPenguinHarnesskimi/kimi-k2Move routes to the new clientWorking13:41
  • lint-passPenguinHarnessdeepseek/deepseek-chatFix the lint errorsWorking4:08
  • docs-passPenguinHarnessopenrouter/qwen/qwen3-coderUpdate the READMEFinished7 min. ago

Usage & cost

Read from each card’s Trace files, so every request belongs to exactly one card. Tokens are exact; subagents’ requests are marked.

AgentTokensCost
coupon-fix1,218,804no price
api-migration963,117no price
docs-pass208,455$0.13
Total2,390,376$0.13
Exact to the picodollar; rounded only on screen.Cards on OpenRouter are priced from its catalogue; a model without a known price shows “no price”.
Under the hood

The real penguin chat, on NeuroSquad’s own server

NeuroSquad starts Penguin’s own chat in a real terminal in the card, attached to the server it runs for its cards. This is everything it adds:

What NeuroSquad adds to the command
❯ penguin chat
--server http://localhost:<the app’s Penguin server>every launchthe app’s own Penguin server, never one started from a card
--agent-id ns_<card id>every launchthe card’s own Penguin agent, rebuilt on every launch
--approve always-askevery launchevery call goes to the card’s hook first
PENGUIN_HOME=<userData>/penguin/dataevery launchthe app’s data root, so the chat finds its server token
PENGUIN_UPDATE_CHECK=offevery launchno update check
--provider <group> --model-id <id>first launchthe card’s model on a fresh session
--resume <session-…>later launchesthe session the stream last reported
pre_tool_use → deny exec_command, input_commandcanvas modethe shell refused; the canvas’s cards instead
pre_tool_use → allow every other calldangerous modedangerous mode, live: no restart

Your ~/.penguin is never written

Your project config is copied into the app’s root on every launch, so cards run on your models and keys; your own root is only read.

No secrets in the card’s environment

The server runs with the app’s clean environment, and provider keys stay in the app’s relay — so a card’s token never reaches Penguin’s shell commands.

Found where you installed it

npm’s package or Penguin’s installer layout, run through its own Node past the penguin.cmd shim.

FAQ

PenguinHarness in NeuroSquad, answered

Give PenguinHarness a canvas

Free for Windows and macOS. Your code, your models and your costs stay on your machine.

Windows 10 / 11 · macOS 12 or newer, Apple silicon and IntelAll 22 agent CLIs