PenguinHarness,with a canvas around it
Run as many Penguin agents as you need, each the real penguin chat in its own card. NeuroSquad runs its own Penguin server and gives every card an agent of its own; Penguin’s status stream tells the card when it works, waits or finishes, and arrows hand it terminals, notes and other agents.
Official sitepenguin.ooo- The real
penguin chat, your own models - Your
~/.penguinis never written - Free, for Windows and macOS
It knows when Penguin is done — and when it’s asking
PenguinHarness is a client and a server. NeuroSquad listens to the server’s own status stream for every card’s session, and a small hook package in each card’s agent decides which tool calls need you.
session_state: runningsignalThe server’s stream reports the card’s session running. The card’s icon spins, the Squad status card counts it, and the toast from its last question is taken down.
pre_tool_usesignalThe card’s pre_tool_use hook leaves a call to Penguin’s own “? Approve this tool call? [Y/n]”. The toast names it — PenguinHarness needs your permission: exec_command: rm -rf dist — until you answer in the card.
session_state: idlesignalThe stream reports idle and the chat is back at its > prompt. The next queued prompt goes out, the journal gets the answer, and a toast says it’s done.
The hook is the policy
The card runs Penguin with --approve always-ask, and its hook answers first: NeuroSquad’s own tools, read_file and subagents run unasked, as in Claude Code; everything else waits on you.
Subagents show up too
A subagent runs in a child session; the stream reports it, so the card stays working while it runs and Usage lists it apart.
No early “finished”
“Finished” waits until the chat is back at its prompt — otherwise a queued prompt pasted a moment too early would be lost.
Three things you’ll do with PenguinHarness
Replicas of the app, played step by step. When Penguin asks, the answer is yours.
Penguin runs the tests through a terminal card and reads the file unasked, stops on its own approval question before editing, and finishes — the journal writes itself.
Everything PenguinHarness can do in NeuroSquad
Measured on PenguinHarness 0.2.13 and on its main branch. Some of it is Penguin’s own — MCP, sessions, its Trace files — and much is built around its server. Where it can’t be done, the card says so.
- 01
Knows its state
Penguin’s own status stream and the card’s hook report every turn and every call that needs you.
- Built by NeuroSquad
Finished and waiting are different
PenguinHarness 0.2.13 has no per-prompt hook, so the card follows the server’s own status stream (
Docs: Finished and waiting are differentsession_staterunning and idle) for its session — every turn, including ones started from the phone. - Built by NeuroSquad
A notification that says what it needs
The toast names the call —
Docs: A notification that says what it needsPenguinHarness needs your permission: write_file: a.txt— and the card goes back to working when you answer it in the card. - Built by NeuroSquad
Stops through its server
Stop, the budget brake and the phone’s Stop call Penguin’s own abort API. Ctrl-C on an idle chat would ask to quit, so Stop never sends it, and the phone’s key row marks Ctrl+C as quitting.
Docs: Stops through its server Picks up where it left off
The stream and the hooks report the session id; the next launch is
Docs: Picks up where it left offpenguin chat --resume <id>. A deleted session starts fresh by itself.The Squad status card
Every AI agent of the workspace on one card, with its status, model and turn time.
Docs: The Squad status card
- 02
Works through the canvas
An arrow from the card is a set of tools, over Penguin’s own MCP client.
Arrows are tools
The card’s agent lists NeuroSquad’s MCP server in its
Docs: Arrows are toolstools.mcpServers. An arrow to a terminal, note, browser or agent gives Penguin those tools.- Built by NeuroSquad
The arrow is the consent
The card’s hook lets NeuroSquad’s own tools,
Docs: The arrow is the consentread_fileand subagents run without a prompt — what Claude Code runs without asking. Everything else waits on you. - Built by NeuroSquad
Tools offered up front
Penguin takes a snapshot of a model context’s tools and ignores list changes. So a card gets every NeuroSquad capability from the start, and the arrows decide at call time.
Docs: Tools offered up front - Built by NeuroSquad
Canvas mode that holds
The hook refuses
Docs: Canvas mode that holdsexec_commandandinput_command— the shell, and Penguin’s only way to the web — before any approval, dangerous mode included. MCP servers from the registry
Installed servers come as a second server with no pre-approval: Penguin asks before every call to it.
Docs: MCP servers from the registry
- 03
Built for long runs
Context, Compact, rate limits, a prompt queue and a journal — handled on the card.
Context meter
The latest request’s tokens from the Trace, against the window Penguin records for the session’s model.
Docs: Context meter- Built by NeuroSquad
Rate limits on the card
When a request ends on a rate or quota error, the card reads it from the Trace and shows the limit.
Docs: Rate limits on the card Prompt queue
Prompts to a busy card wait and go out after the turn, one by one; long and multi-line prompts arrive as one message.
Docs: Prompt queueJournal and hand-off
The last answer is read from Penguin’s Trace files: into a connected note, a hand-off,
Docs: Journal and hand-offagent_wait, Telegram and custom cards.
- 04
Works in a team
An agent per card, subagents, live dangerous mode and a switch to another CLI.
- Built by NeuroSquad
An agent per card, its own server
NeuroSquad runs its own Penguin server for its own data root and stops it at quit. Each card is a Penguin agent of its own, with its own hooks, MCP servers and instructions; your projects and models are copied in, never changed.
Docs: An agent per card, its own server Subagents, counted apart
A subagent runs in a child session with its own Trace. The card stays working while it runs, and Usage, Run Stats and Agent Pulse mark its requests as a subagent’s.
Docs: Subagents, counted apart- Built by NeuroSquad
Dangerous mode, live
No restart: while the card’s switch is on, its hook allows every call; turn it off and the next call asks again. Canvas mode still wins.
Docs: Dangerous mode, live Model and CLI switch
A lead agent can change the model with
Docs: Model and CLI switchagent_set_model— in place on newer Penguin; on 0.2.13 a new session starts with Penguin’s own pointer to the old one — or move the conversation to Claude Code and back.Isolated worktrees
Give an agent its own git worktree, so parallel agents never edit the same files.
Docs: Isolated worktrees
- 05
Cost, control and access
What it spends, what it may spend, which model it runs on, and how you reach it.
Exact usage
Read from Penguin’s Trace: one record per request, cache reads apart, subagents marked. A model with no known price shows “no price”, never $0.
Docs: Exact usageBudget brake
Past the workspace limit each Penguin card is stopped through the server’s abort API — the chat stays open — and automatic prompts wait.
Docs: Budget brake- Built by NeuroSquad
OpenRouter models
A provider group of the card’s own goes through NeuroSquad’s local relay, which adds the key and the attribution; the file holds only a placeholder. Your own OpenRouter group is routed the same way.
Docs: OpenRouter models - Built by NeuroSquad
Your own servers
Each of your providers becomes a group behind the same relay — OpenAI Chat Completions or Anthropic Messages — with no key on disk.
Docs: Your own servers From your phone
Scan a QR code and the whole canvas opens in the phone’s browser: read Penguin’s chat, answer y or n, send the next prompt.
Docs: From your phone
- 06
Good to know
What Penguin does differently, and what can’t be done.
- Not possible
The card can’t pick the session id
Penguin always creates its own ids. NeuroSquad catches the id from the stream and resumes with it, which works the same in practice.
Docs: The card can’t pick the session id - Not possible
No cache-write or reasoning split
Penguin’s Trace folds cache creation into the input and reasoning into the output, so those columns read “not reported”. On Claude models the price is a little low for it.
Docs: No cache-write or reasoning split - Not possible
No token saver
Penguin’s
Docs: No token saverpre_tool_usehook can only allow or deny a call, not rewrite it — so RTK can’t work there. Plugin doors need a newer Penguin
Caveman, Memory, Context7 and Graphify add text through a per-prompt hook that Penguin runs on every prompt only after 0.2.13. On 0.2.13 their tools still work by arrow.
Docs: Plugin doors need a newer PenguinThis computer only
The Penguin server would have to run inside a WSL or SSH target, so Penguin cards run on this computer only for now.
Docs: This computer only
Who’s waiting, and what it used
Two cards that answer the questions you ask most when several Penguin agents run at once.
The Squad status card
Every agent of the workspace with its status, model and turn time. The ones waiting on you come first, with the call they want to make. Tap a status to filter.
coupon-fixPenguinHarnessdeepseek/deepseek-chatPenguinHarness needs your permission: exec_command: rm -rf distNeeds you2:34
api-migrationPenguinHarnesskimi/kimi-k2Move routes to the new clientWorking13:41
lint-passPenguinHarnessdeepseek/deepseek-chatFix the lint errorsWorking4:08
docs-passPenguinHarnessopenrouter/qwen/qwen3-coderUpdate the READMEFinished7 min. ago
Usage & cost
Read from each card’s Trace files, so every request belongs to exactly one card. Tokens are exact; subagents’ requests are marked.
| Agent | Tokens | Cost | ||
|---|---|---|---|---|
| coupon-fix | 1,218,804 | no price | ||
| api-migration | 963,117 | no price | ||
| docs-pass | 208,455 | $0.13 | ||
| Total | 89 | 2,390,376 | $0.13 |
The real penguin chat, on NeuroSquad’s own server
NeuroSquad starts Penguin’s own chat in a real terminal in the card, attached to the server it runs for its cards. This is everything it adds:
Your ~/.penguin is never written
Your project config is copied into the app’s root on every launch, so cards run on your models and keys; your own root is only read.
No secrets in the card’s environment
The server runs with the app’s clean environment, and provider keys stay in the app’s relay — so a card’s token never reaches Penguin’s shell commands.
Found where you installed it
npm’s package or Penguin’s installer layout, run through its own Node past the penguin.cmd shim.
PenguinHarness in NeuroSquad, answered
Give PenguinHarness a canvas
Free for Windows and macOS. Your code, your models and your costs stay on your machine.