A coding agent runs in one tab and a research chat in the next. Each one is capable, and you are the person carrying context between them. This guide shows how a coordinator assistant splits the work, passes it on and keeps you in charge of what goes live.

Five chat windows make you the router

A typical week for a small business owner who uses AI seriously looks like this. Cursor is fixing the website while Codex writes a script for the invoice export. Every session starts with you pasting the same background again.

When one agent finishes, you copy its result into the next window. The agents are fast; the slow part is you, moving text and remembering which version is current.

The Grok Bot documentation describes the alternative directly. Its bots "can run in parallel, message each other, share context in group chats, and pass ownership of a task, so you aren't the router between tools" (Grok Bot overview).

Engineers have a name for this shape: orchestrator-workers. Anthropic describes a central model that breaks a task down, delegates parts to worker models and combines their results. It names coding as a good fit, because the number of files a change touches depends on the task and can't be planned in advance (Anthropic, Building effective agents).

Why a coordinator works better than many windows

The first reason is context. A specialist agent works best with a narrow brief and a clean workspace. Claude Code's documentation says each subagent starts with a fresh, isolated context window and doesn't see the main conversation history; it works from a delegation message and returns a summary (Claude Code subagents).

That keeps the main conversation readable, and it puts real weight on the handoff. The coordinator holds the plan, and the plan has to live somewhere more durable than a chat. In Anthropic's research system, the lead agent saves its plan to memory, because a context that grows past 200,000 tokens gets truncated (Anthropic, multi-agent research system).

The second reason is isolation. Cursor cloud agents run in isolated virtual machines, clone the repository, work on a separate branch and push changes back for handoff (Cursor Cloud Agents). Codex Cloud gives each task its own workspace, and you review the changes before you commit or open a pull request (Codex Cloud).

A coordinator uses both properties and reports back to you in one place. An engineer on the Grok Bot team describes bots that start Cursor cloud agents, read their transcripts, check the proof attached to pull requests and send follow-ups or interrupt a run (Grok Bot for Engineering).

How the work splits

In practice the split follows the kind of thinking each step needs:

  • Planning. The coordinator turns a request into tasks, each with a finish line and the proof it expects back.
  • Coding. A cloud coding agent such as Cursor, Codex or Claude Code works on its own branch.
  • Review. A second agent reads the diff. Codex can review pull requests automatically or when someone mentions it in a comment (Codex code review on GitHub). You read the summary and the proof.
  • Research. A separate agent gathers sources and returns short findings with links, so the coordinator's context stays clean.

The Grok Bot engineering guide adds that its bots "perform best when focused on a single domain" (Grok Bot for Engineering).

Set it up in five steps

Step 1: Write the standing rules into a file

Rules typed into a chat disappear with the chat, so put them in the repository. Codex reads AGENTS.md files before doing any work, layering global guidance with project rules (Codex AGENTS.md guide). Write down which branch to use and which folders stay untouched.

Step 2: Give every task an objective, a format and a boundary

Anthropic found that each subagent needs an objective, an output format, guidance on tools and sources, and clear task boundaries. Without detailed task descriptions, its agents duplicated work or left gaps (Anthropic, multi-agent research system). Add the proof you expect, such as a screenshot of the changed page or the test output.

Step 3: Hand off through branches and files

A branch is a handoff that survives a closed laptop. GitHub's Copilot cloud agent can push to one branch only, a new copilot/ branch or the existing pull request branch, and it stays subject to branch protection (GitHub, risks and mitigations for Copilot cloud agent).

When two agents could touch the same files, separate them. Claude Code's documentation recommends worktrees, a separate checkout per session, and says agent teams need the work partitioned so each teammate owns different files (Claude Code, run agents in parallel). Keep one task list the coordinator updates: task, agent, branch, status.

Step 4: Put a person on every irreversible step

The Grok Bot documentation lists sending messages, publishing, purchases, deleting data, changing permissions and production changes as places for explicit boundaries. It also notes that "an approval controls the proposed action. It does not reverse work already completed" (Grok Bot approvals, security and privacy).

GitHub builds the same idea into its coding agent. Its draft pull requests must be reviewed and merged by a human, the agent can't approve or merge them, and by default workflows wait until someone with write access approves the run (GitHub, risks and mitigations). We use one plain rule: an agent may prepare anything on a branch or as a draft, and deploying, spending and publishing wait for my approval.

Step 5: Keep a log someone can read later

Cursor cloud agents attach screenshots, videos and logs to the pull request (Cursor Cloud Agents). GitHub links each agent commit to its session log (GitHub, risks and mitigations). The Grok Bot least-privilege checklist ends with "preserve source links and an action log for important decisions" (Grok Bot approvals, security and privacy).

A log also answers the ownership question when someone leaves. The same reasoning applies to agents running inside Microsoft 365: every agent needs a named person who answers for it.

Where it breaks

Context loss. A worker knows only what the handoff gives it. The opposite risk exists too: in the OpenAI Agents SDK, the receiving agent sees the entire previous conversation by default unless you filter it, and nested history "does not redact sensitive data" (OpenAI Agents SDK, handoffs). Decide what each handoff carries. For consequential decisions, the Grok Bot documentation advises checking the current source instead of relying on memory (Grok Bot overview).

Duplicated work. Anthropic's early agents spawned 50 subagents for simple queries and distracted each other with excessive updates (Anthropic, multi-agent research system).

Cost. In Anthropic's data, agents used about four times more tokens than chat, and multi-agent systems about fifteen times more. Anthropic also notes that most coding tasks have fewer truly parallel parts than research (Anthropic, multi-agent research system). Set a usage limit before you add agents; the approach from setting Copilot credit limits first carries over.

Over-automation. Anthropic's advice is to find the simplest solution and add agents only when they demonstrably help, because autonomy brings higher cost and compounding errors (Anthropic, Building effective agents). Automated review has limits as well. Grok Bot describes its Auto Review as model-based, a complement to least privilege and explicit approval boundaries rather than a replacement (Grok Bot approvals, security and privacy).

What not to do

  • Don't give any agent push rights to the main branch or to production.
  • Don't paste passwords or one-time codes into a chat. Enter them yourself in the tool that asks.
  • Don't treat separate bots as separate security zones. Grok Bot's bots share one cloud computer, files and logins included, and its documentation says not to use separate bots as a security boundary (Grok Bot approvals, security and privacy).
  • Don't start four agents for a change one agent can finish in ten minutes.
  • Don't merge because the automated review came back clean. Read the summary and look at the proof.

When you can run this yourself

You can start this on your own if you already use one coding agent and keep your code in a repository with branches. Begin small:

  • Pick one recurring task with a clear finish line.
  • Write the standing rules into AGENTS.md or the equivalent file for your tool.
  • Let the coordinator start one specialist on one branch and require proof back.
  • Protect the main branch and keep merges, deploys, spending and publishing behind your approval.
  • Keep a task list and an action log, then review both after the first two weeks.
  • Add a second specialist only when the first one waits on work it can't do.

Bring in help when agents need access to customer data, payment systems or production servers, or when nobody in the business can read a diff. A coordinator works like a good site foreman: it knows which crew is on which wall, and the drawings stay in the site office where anyone can check them. If you want that set up and maintained alongside your other systems, it fits into ongoing tech partnership work.