Quickstart guide
Getting started with agents
Most companies do not fail at agents because the models are not good enough. They fail because nobody decided which work the agent was for, agreed what success looked like, or could see what the agent actually did.
Five steps, in order, from zero to agents doing real work with a number attached.
Step 01
Start with the why, then the work
Start with the outcome, not the tool. What are you trying to achieve — cut time from request to shipped, clear a backlog, lower the cost of something you do at volume? Name the why first. Then find the work that serves it: repetitive, well documented, text or code heavy, with a clear definition of done. Set two or three metrics you can measure today — outcome, not activity — before anything runs.
Worked example
“I spend about eight hours a week triaging inbound support tickets — reading, tagging, and routing each one.”
- How I'd start
- The agent drafts the tag and routing on every ticket; a person approves before it applies. As overrides fall, approvals get lighter. The number to move: tickets triaged per hour, measured two weeks before and after.
Where Spirehawk fits
Baselines are hard because the data is scattered. Spirehawk captures every agent session and ties it to real work — so the numbers you set in week one are the ones you report in month three.
Step 02
Choose the agent that fits the work
Do not run a vendor beauty contest. Choose by the shape of the work, and expect more than one agent — one agent for everything is mediocre at most of it.
Terminal and codebase agents (Claude Code, Codex): repo work — migrations, test coverage, bug triage. Highest return, and easiest to verify because tests either pass or they do not.
Desktop and knowledge-work agents (Claude Desktop): research, drafting, analysis. The best first step for teams outside engineering.
Embedded and API agents: agents inside your own product. Highest ceiling, highest effort — third, not first.
Worked example
“Ticket triage is my workflow. Which agent runs it — and how do I avoid lock-in?”
- The pick
- A desktop, knowledge-work agent — triage is reading and classifying text, not repo work. Run two candidates on the same flow through the gateway and compare them on identical numbers.
Where Spirehawk fits
Spirehawk is not another agent — it sits at the gateway, so any gateway-configurable agent reports into one place. Run two on one workflow, compare them on identical metrics, and swap without losing your history.
Step 03
Run one scoped pilot
One workflow, one team, two to four weeks, a human approving output before it reaches anything real. Run five pilots at once and you will not know what caused the numbers.
Scope it tight. “Agent drafts first-pass triage on inbound tickets” is a pilot. “Agent improves support” is a wish.
Count the interventions, log the failures. A falling rate over four weeks is your strongest evidence it will scale; flat means the workflow is wrong, not the agent. The failure log becomes your Step 4 guardrails.
Set a kill criterion up front. If it has not beaten the baseline by week four, you stop. Knowing early is worth the spend.
Worked example
“One support team, four weeks: the agent drafts triage, a person approves before anything is applied.”
- What I'd watch
- The human override rate — falling over four weeks is the strongest sign this will scale. End on a one-page result against the baseline and a clear go or no-go.
Where Spirehawk fits
In a pilot you need to see the work as it happens. Spirehawk's rail shows every live session and logs every approval and denial — so the intervention rate is a number you can pull, not a memory you chase.
Step 04
Scale it to the team
Most companies skip this step, and it decides whether agents stick. A pilot proves the technology works. It does not prove your organisation will use it.
Champions, not mandates. Give each pilot user two or three colleagues to bring along. Peer-to-peer beats a training deck.
Teach the workflow, not the tool. Twenty minutes on the prompt pattern that works here, where it fails, and when to take over — not a lecture on transformers.
Write the guardrails down. Turn the Step 3 failure log into a one-page policy: what agents may touch, what needs approval, what they never see.
Name an owner accountable for adoption and reporting the Step 1 metrics — and budget a slower fortnight while people ramp.
Worked example
“The pilot beat the baseline. Now the whole support org needs it — without the risk scaling with it.”
- The guardrails
- Turn the pilot's failure log into a one-page policy: what the agent may tag, what needs approval, what it never sees. Each pilot user brings two colleagues along — peer-to-peer beats a training deck.
Where Spirehawk fits
At scale, visibility is the control system. Spirehawk's Govern layer enforces guardrails across every agent and person — risky actions pause for approval, new joiners inherit the rules, and adoption scales without your risk surface scaling with it.
Step 05
Prove the value, and keep proving it
An agent programme that cannot show its return gets cut in the first budget review, even when it works — the failure is rarely the agents, it is that nobody produced the evidence. What did we get back for the spend, in the CFO's currency? Can we prove to an auditor that agents are governed? Neither is answerable from a chat log. This is what Spirehawk is for.
Observe. Every session from every agent on one live rail, plus searchable history — running in minutes at the gateway, without instrumenting a single agent.
Measure. Sessions attributed to real work — report cycle time, throughput and cost per outcome against your Step 1 baseline, model spend split by team and workflow.
Govern. Scope what agents may touch, approve or deny risky actions before they run, and keep the log the auditors will ask for.
Worked example
“Finance is asking what we got back for the model spend. I need an answer I can defend.”
- The number
- Cost per triaged ticket against the human baseline, plus time to first response, month over month — backed by a searchable log you can show an auditor.
Where Spirehawk fits
Every answer here — the live rail, cost tied to outcome, the audit log — comes from the one gateway, so the CFO gets a number and the auditor gets the trail from the same place.
Start with Step 1 today
Instrument first, then pilot — the Observe layer runs in minutes at the gateway, so the numbers you set in week one are the ones you report in month three.