← Back to side projects

Mission Control

A private dashboard and chat for managing my team of AI agents.

Problem

I had a team of AI agents building things for me, and I needed a way to track, evaluate and manage their work.

Solution

A private dashboard and chat for the team, with a tap-to-approve card for anything risky.

So far

59 merged pull requests in its first week, backed by an automated test suite.

The tour

  1. See everything at a glance

    One screen shows service status, server health, and what is waiting on me or new. I never have to check each agent one by one.

    Mission Control overview on a desktop: an approvals strip reading 1 waiting, the inbox, today's research digest, the agents list and the active projects
  2. Talk to the team. Approve with a tap.

    Risky actions (deploys, merges, public posts) stop and wait for a deliberate hold-to-approve on a card. A typed "yes" doesn't count, because text can be quoted or faked.

    Chat on a phone: a pixel-art mascot, and a high-risk production deploy approval card with Decline and Hold to approve buttons
  3. Know the team

    Like a team roster: what each agent is for, which tools it may use, and whether it is working right now. Clear roles are what make delegation safe.

    The Agents page: a grid of agent cards such as orchestrator, architect, designer, fullstack-dev, code-reviewer, qa-tester, devops and tech-writer, each with its model, tools and run count
  4. One list of ideas

    Ideas don't get lost in chat. Picking one up hands it to the orchestrator with its context.

    The Backlog page: an ordered list of three items, then items grouped by feature, each with a Pick up button
  5. See how the work connects

    Agents, projects, notes and runs as one map, so I can see where work clusters and what depends on what.

    The Graph page: agents, projects, notes and runs drawn as a rotating sphere of linked nodes, coloured by project

Screens are from my real setup, with nothing private shown.

How it's built

Where AI agents are going (and what it means for PMs)

Agents went from single tools to teams to platforms, and the next step is full autonomy, which is held back by trust more than by capability.

On my own time, I run a team of about 15 AI agents on a small home server. An orchestrator agent takes my requests and hands them to specialists: an architect, a designer, a developer, a code reviewer, QA, devops, writers, a small video studio and a researcher. I steer it all from a dashboard I built called Mission Control. In its first week, starting September 30, the team merged 79 pull requests across my side projects.

Running it has shown me where this is heading. I see four steps so far. The fifth matters most, and it is a management problem.

  1. Agents with context and tools
  2. Agent teams and orchestration
  3. Evaluation and the Agent OS
  4. Full autonomy
  5. Trust

Trust is the open stage.

Agents with context and tools

The first wave was about making one agent useful. Give a model the right information and the right tools and it stops guessing. Anthropic's Building effective agents drew the early line between agents that direct themselves and workflows, "systems where LLMs and tools are orchestrated through predefined code paths."

The conversation then moved from prompts to what Anthropic calls context engineering: "the set of strategies for curating and maintaining the optimal set of tokens (information) during LLM inference."

In my setup, each agent has a written definition, a short list of tools and a lessons file it reads before every job. Most of my early gains came from editing those files.

Agent teams and orchestration

One agent hits a ceiling, so the next step was teams. Anthropic's multi-agent research system beat a single agent by 90.2% on its internal research eval, at about 15 times the tokens of a normal chat. At Cursor, Michael Truell writes that 35% of the company's internally merged PRs are created by autonomous cloud agents.

My orchestrator never writes code itself. A feature goes from architect to designer to developer to reviewer to QA, and then to deploy. It looks like a small product team because it is one. As the Anthropic post warns, "Agents are stateful and errors compound."

Evaluation, and the rise of the Agent OS

Once you have a team, the bottleneck moves. Output gets cheap. Knowing whether it is any good stays expensive. Anthropic's guide to agent evals says that without evals, teams "can't distinguish real regressions from noise" and are "flying blind."

That is why people are building what I call an Agent OS. There is no agreed definition, so here is mine: the layer that runs, evaluates and governs agents, tracking who may do what, what happened, whether it was good, and who signed off. Vendors call their versions an agent platform or a control plane. Writing about Google Cloud Next 2026, Forrester described Google's pitch as "building, securing, running, governing, and observing agents in the enterprise" and concluded: "The agentic stack has consolidated into a single agent platform."

Mission Control is my small, personal version. It shows what each agent is doing, holds the approval cards and links to every plan, review and QA report. Every reviewer ends with a check ledger: each check it planned, marked pass, fail or not run. A check that was skipped can't quietly disappear into "looks good."

Full autonomy: the user gets the end result

The next step is that you stop managing the steps and receive the result. The signs are already here. Anthropic measured that the longest Claude Code turns (the 99.9th percentile) grew from under 25 minutes to over 45 minutes between October 2025 and January 2026. Truell describes agents that hand back "logs, video recordings, and live previews rather than diffs."

We are already seeing this play out with skills like the gauntlet loop. The name and core idea are public: Matt Shumer popularized it, and the rule at its heart is "never let the builder grade its own homework." Builders produce the work, a separate critic that never sees their reasoning judges it against a concrete bar, and the loop repeats.

I use my own version. For my portfolio home page, three rival builder agents each made a design. Each round, a fresh, blind critic agent judged them against a written brief and three reference sites, with measured evidence: sideways scrolling on a phone, tap target sizes, contrast, animations left running. Hard cap of three rounds. Nothing passed in rounds one or two. Round three failed on three cheap defects, one of them a caption that described an image as something it wasn't. After the fixes, a targeted re-check passed. I did not pick the winner by eye, and the critic caught a false caption I would probably have missed.

Reverse-engineering skills point the same way. Reverse Engineer Anything (REA) is an open-source MCP server that lets an agent take apart an existing app with classic decompiling tools, write up how each feature works, and keep looping until it has rebuilt the app. In his test of it, Nick Saraev makes the bigger point: you can often skip the library and simply ask an agent to work out how a feature you like works and rebuild it. Either way, you describe the result you want and the agent finds its own path there.

The real blocker is trust

The question is: do you trust the way your team is built enough to let it go off on its own?

The best answer I have heard comes from Lauren Tan, an engineer on Grok Bot at SpaceX AI who previously worked at Cursor and on the React Compiler team. In a recorded interview she makes the management parallel directly:

for me, the parallel is like with management. If I'm an engineering manager of a team [...] and I don't trust them then the mode of operation I'm going to be in is going to be like micromanagement.

Anyone who has hired someone knows this curve. In week one you read everything they write. After a month you check outcomes. After a quarter you read the weekly update and step in when something looks off. Trust comes from a track record and from the systems around the person.

Tan says the same about scale: "you can't go to a 100 agents, like spawn 100 agents, when you don't even trust the output of one agent." Her answer is verification, which she calls "the most important skill that you should have in your toolbox when you work with agents": an agent that can run and test its own work. Without it, "you are the verifier. You're the bottleneck."

Her path is a ladder: verification locally, then trust locally, then cloud agents picking up bug reports and returning PRs, then auto-merging, with her reviewing changes after they land. She describes waking up to "like 20 PRs landed and I just reviewed them on main like they were already landed and they were good."

Anthropic's autonomy data shows the same curve: new users fully auto-approve about 20% of sessions, over 40% by 750 sessions. The flip side is fatigue. Anthropic reports that users approve 93% of permission prompts, and warns about people who "stop paying close attention to what they're approving."

My setup is my version of that ladder. Anything risky waits for my tap on an approval card in Mission Control: merges, deploys, public posts, spending. A "yes" typed in chat doesn't count, because text can be quoted, pasted or faked, and a tap on a specific card can't. Then I noticed I was tapping routine merge cards within seconds. That was the 93% problem in miniature: the card had stopped being a review. So on October 9 I started a one-week autonomy trial. Changes that pass code review and QA now ship in a release train twice a day, under one card that lists each item, its verdicts and how to roll it back. Public, destructive and money decisions still get their own card.

Tan has one more idea I've taken on. If a human has to catch the same problem in review twice, treat that as "a code smell" and ask "how do I turn this into a lint rule? How do I turn this into a CI failure?" In my team, every repeated mistake becomes a one-line lesson or a check that agents read before the next job. That is how trust compounds. Autonomy is something you grant one rung at a time, after the work has earned it.

What this means for PMs

Truell's description of the new human role reads like a PM job description: "The human role shifts from guiding each line of code to defining the problem and setting review criteria."

Marty Cagan warns about the trap. In The AI Productivity Paradox he argues that AI speeds up teams whose model is "designed to deliver output, rather than outcomes," so they ship faster without getting better results. He quotes Chip Huyen: "AI makes building easier, but the hardest part remains knowing what to build."

So the work shifts. Defining the problem matters more, because agents will happily build the wrong thing quickly. Writing review criteria becomes core PM work: evals are acceptance criteria that a machine can check, and Teresa Torres already writes about evals as part of discovery. And PMs will decide how much autonomy each workflow gets, which is a product decision about risk, the same call a manager makes about a new hire.

Agents will keep getting more capable. How much of it you can use depends on how well you have defined what good looks like, and that has always been the PM's job.