Agent runs are not timesheets
An agent works while nobody is watching. That work still needs a record with an owner, a scope, a cost, and a reviewer, and an hours grid cannot hold one.

An agent ran for forty minutes at three in the morning. It refactored a module, opened a pull request, and stopped. Nobody sat through it. In the morning a developer read the diff, corrected two things, and shipped it.
Now write that down. A timesheet wants an hour and a person, and there is no person to name for those forty minutes. An AI dashboard offers a token count, which is a measure of consumption rather than of work delivered. Neither is a record you would put in front of the client who paid for the outcome.
This is the third mode of work, and it is the one the industry has no format for. Human work has timesheets. Human-directed AI work fits an existing entry with an involvement level attached. Autonomous runs fit neither, because the unit is not an hour of attention. It is a run.
Example record
One delivery, three distinct facts
Owner
Maya Chen
Responsible for the final delivery
AI contribution
Generated
Prepared the first test-suite draft
Review
Verified
Assumptions checked and brittle tests corrected
Why a run is not an hour
An hour on a timesheet carries an implicit promise: a person was engaged, judgment was applied, and the elapsed time is a fair proxy for the effort spent. Every downstream use of that number, from billing to capacity planning to estimating the next project, leans on that promise.
A run breaks the proxy. Forty minutes of agent execution is not forty minutes of anyone's attention, and it is not forty minutes of comparable effort. Logging it as an hour of human time invents work that nobody did. Discarding it loses the fact that a deliverable was produced and that it cost real money to produce.
So a run needs its own row, with fields that describe what a run actually is.
What a run record has to contain
Seven fields cover the questions that come up later, in review, in billing, and in a client conversation about how the work was done.
| Field | What it answers | Why it matters later |
|---|---|---|
| Initiator | Which person or schedule started it | Every run traces to a human decision |
| Scope | Project, workstream, deliverable | Puts the run in the delivery record, not a generic AI bucket |
| Duration | Wall-clock execution time | Distinguishes a long job from an expensive one |
| Cost | Model spend for the run | The only honest input to cost per outcome |
| Output | What it produced or changed | Turns the run from activity into delivery |
| Review | Who checked it, and when | The accountability link, and the field most often missing |
| Result | Accepted, corrected, or discarded | Tells you whether the run was worth repeating |
Review and result are the two that teams skip, and they are the two that make the record worth keeping. A run nobody reviewed is not finished work. A run whose output was discarded still cost money, and a month of discarded runs is a signal worth acting on.
The human stays the headline
An agent cannot own a deliverable. It cannot be accountable to a client, cannot be asked what it was thinking, and cannot answer for a mistake six months later. Whatever the automation did, a person accepted the result and put their name on it.
Practically, that means every run record names two people: the one who initiated it and the one who accepted the output. They are often the same person, and writing both down anyway is what makes the exceptions visible. A run initiated by a schedule and accepted by nobody is exactly the case you want to be able to find.
Where the run attaches
Attach runs to the same workstream structure that human time uses. If the agent refactored the payments module, the run belongs to the same workstream as the human hours spent reviewing that refactor, not to an AI category sitting off to one side.
This is the difference between a record you can report on and a pile of logs. When runs and hours share a spine, you can answer the questions people actually ask: what did this deliverable cost in human time and machine cost, how much of it was reviewed, and how does that compare with the same work three months ago.
A separate AI category cannot answer any of those. It can only tell you how much AI you used, which is the least interesting thing about it.
What not to record
The temptation with autonomous work is to capture everything, because everything is capturable. Resist it, for two reasons. Prompt contents and intermediate reasoning are frequently confidential, sometimes the client's confidential material rather than yours. And a record nobody can read is a record nobody reviews.
- Prompt text and intermediate steps: keep a pointer to the run if your tooling stores one, not a transcript in the delivery record.
- Keystroke or screen activity: this measures the person, not the work, and it changes behaviour in ways that make every other number less honest.
- Per-message token counts: aggregate cost per run is the number a client conversation can use.
The test is whether a field would change a decision. Cost per run changes model routing. A transcript changes nothing until something goes wrong, and then you want the trace in your engineering tooling, not in the deliverable's public record.
What a run record catches that a timesheet cannot
Three patterns only become visible once runs have their own rows, and each one costs money while it stays invisible.
The first is silent cost growth. A workflow that cost eight dollars a day in March and forty in July did not announce the change. Nobody approved it, because nobody was looking at cost per outcome, only at a monthly bill that grew slowly enough to look like adoption.
The second is the discard rate. When a third of runs produce output that gets thrown away, the tooling is being pointed at the wrong problem. Without a result field this looks like productivity, because the runs happened and the tokens were spent.
The third is review debt. Runs accumulate faster than people can check them, and the gap widens quietly until a client finds something nobody read. A record with a review field turns that into a number you can watch instead of an incident you discover.
What it does to estimating
Estimates are built from what comparable work took last time. Once part of the work is machine execution, a single hours figure stops being comparable, because two projects with identical hours may have involved very different amounts of automated work and very different amounts of review.
Keeping runs and hours separate but attached to the same workstream fixes the comparison. The next estimate can say that similar work took nine human hours, roughly twenty runs, and about sixty dollars of model spend, and that the review step took longer than the drafting did. That is a forecast. A blended number is a guess wearing a forecast's clothes.
What the client sees
Most clients never need the run-level detail, and offering it unprompted reads as noise. What they need is the shape of the work: this deliverable took eleven hours of human time, involved agent execution that cost forty dollars, and was reviewed by a named senior engineer before it shipped.
That paragraph is only writable if the runs were recorded properly in the first place. The record is not for the client. It is what makes the sentence you tell the client true.
Where to start
Pick the one automated workflow that already runs regularly, and give it the seven fields. Not every agent, not every tool, and not a governance programme. One workflow, recorded properly for a month, will teach you more about what your record format needs than a policy document written in advance.
Then look at what the month tells you. If most runs were accepted without correction, the review step can get lighter. If a third were discarded, something upstream is wrong and you now have the evidence to name it.