Report no.
NS-2607-01
Example report · Northstar Studio is an invented firm
AI leverage report
Northstar Studio
Software development · 12 people · issued 31 July 2026
1.86×
Modelled effective output per human hour
- Team
- 12 people
- Rate
- €120 / h
- Maturity
- Rolling out
- Issued
- 31 Jul 2026
Modelled value · effective output per human hour
1.86×
Stated range 1.26× to 2.42× across scenarios
- Research conservative
- Broad field evidence, applied cautiously
- 1.26×
- Northstar's model (selected)
- The selected work mix at rolling-out maturity
- 1.86×
- Task frontier
- A bounded upside case, not a forecast
- 2.42×
| Capacity account | Hours / year | At €120 / hourRevenue that stops appearing on invoices |
|---|---|---|
| Baseline capacity12 people × 1,400 hours | 16,800 | €2,016,000 |
| Human hours in the planHours the team works | 9,038 | €1,084,608 |
| Capacity returnedHours to redeploy | 7,762 | €931,392 |
The same mix, at every maturity
Nothing below changes the work mix. Only how well the team runs what it already has.
| Maturity | Leverage | Capacity returned | At €120 / hour |
|---|---|---|---|
| Backfiring | 0.75× | -5,456 h | -€654,720 |
| Piloting | 1.60× | 6,320 h | €758,400 |
| Rolling outNorthstar | 1.86× | 7,762 h | €931,392 |
| Embedded | 1.95× | 8,195 h | €983,378 |
Method and conditions
Effective output divided by human hours, from the same weighted task-time model as the calculator. This is modelled capacity, not an observed team result. Recoverable only if the freed hours are sold to new work. Unsold, they leave as a smaller invoice rather than a bigger margin.
01 · Drivers
Implementation carries most of the estimate.
Implementation is 32% of delivery time and returns 24.0 of the points below. The three areas with no AI applied are 40% of the time between them, and that is what caps the ratio at 2.50×.
Discovery and scoping
Untouched
0.0 pts
Design and UX
Untouched
0.0 pts
Implementation
4.0× applied
24.0 pts
Testing and QA
5.0× applied
11.2 pts
Code review
4.0× applied
6.0 pts
Documentation
6.0× applied
5.0 pts
Client comms and PM
Untouched
0.0 pts
| Work area | Share before AI | Multiplier | Share after |
|---|---|---|---|
| Discovery and scoping | 10% | 1.0× | 10.0% |
| Design and UX | 12% | 1.0× | 12.0% |
| Implementation | 32% | 4.0× | 8.0% |
| Testing and QA | 14% | 5.0× | 2.8% |
| Code review | 8% | 4.0× | 2.0% |
| Documentation | 6% | 6.0× | 1.0% |
| Client comms and PM | 18% | 1.0× | 18.0% |
| The year, after the multipliers | 100% | 1.86× | 53.8% |
02 · Industry evidence
Real gains, with real scope limits.
A pooled field trial lands at 1.26×, a bounded lab task at 2.26×, and practitioners report speed gains larger than the value they delivered.
Northstar's model1.86×
| Result | Finding | Study design | Source |
|---|---|---|---|
| ≈1.26× | Broad field evidencePooled field experiments found 26.08% more completed tasks across 4,867 developers. | Observed · broad workplace setting | Read ↗ |
| ≈2.26× | Bounded implementation taskParticipants completed a controlled greenfield HTTP-server task 55.8% faster. | Experiment · bounded task | Read ↗ |
| 3× / 1.4–2× | Speed versus delivered valueTechnical workers reported larger gains in speed than in the value of work completed. | Self-reported · convenience sample | Read ↗ |
| Mixed | Mature repository workExperienced maintainers can lose time when generated work creates verification and integration overhead. | Observed downside case | Read ↗ |
Interpretation
The modelled range is a portfolio hypothesis about the mix of work. Any single ticket can land anywhere inside it.
Once review, rework and downstream delivery are counted, only some of the work returns capacity at all.
03 · Opportunities
Practices worth testing first.
A skill here is a written, reusable instruction set the team's agents run, so one practice can be versioned and measured. Each of the three has a failure mode that decides whether its gain survives review.
01 · Implementation
Turn repository context into a reusable change workflow.
Try
Require a code-path survey, stated constraints, project checks and an explicit residual-risk note before a change is considered complete.
Tooling
- Claude Code
- OpenAI Codex
- Cursor
Watch for
Fast generation followed by long integration, review or rework.
- Share of delivery
- 32%
- Multiplier applied
- 4.0×
- Points returned
- 24.0
- Share after
- 8.0%
Evidence for this one · Bounded implementation task · ≈2.26×
The closest published anchor is a greenfield HTTP server, which is the easiest possible case. Northstar's repository is not greenfield, so the METR downside case applies here too.
How WhoWorked measures it
WhoWorked keeps the approved skill, records actual use and compares its cost with human time on the same workstream.
Suggested skill
$repository-change03 · Opportunities, continued
Generated tests can look finished and not be.
02 · Testing and QA
Make the test boundary explicit before generating tests.
Try
Name the changed behavior, affected boundaries and regression risks first; then ask the agent for the smallest useful verification set.
Tooling
- Codex
- Claude Code
- GitHub Copilot
Watch for
More tests with no increase in confidence, or generated tests that merely restate implementation details.
- Share of delivery
- 14%
- Multiplier applied
- 5.0×
- Points returned
- 11.2
- Share after
- 2.8%
Evidence for this one · Broad field evidence · ≈1.26×
No published result measures generated tests on their own. The pooled field figure is the safest reference, and it counts completed tasks rather than confidence in them.
How WhoWorked measures it
WhoWorked shows whether the approved testing workflow is used, where it fails and whether teams drift into shadow variants.
Suggested skill
$test-boundary-analysis03 · Opportunities, continued
Documentation is the cheapest one to start with.
03 · Documentation
Generate the delivery note from the work that happened.
Try
Explain what changed, why, verification performed, operational consequences and what the next person needs to know.
Tooling
- Claude Code
- OpenAI Codex
- GitHub Copilot
Watch for
Polished prose that is detached from the shipped change or hides unresolved risk.
- Share of delivery
- 6%
- Multiplier applied
- 6.0×
- Points returned
- 5.0
- Share after
- 1.0%
Evidence for this one · Speed versus delivered value · 3× / 1.4–2×
This is where self-reported speed runs furthest ahead of delivered value. Prose is fast to generate and slow to verify against the change it describes.
How WhoWorked measures it
WhoWorked tracks which documentation practice survives, identifies private variants and surfaces library skills the team has abandoned.
Suggested skill
$delivery-note04 · Measurement
Four questions that turn an estimate into evidence.
Skill runs, error rates, AI cost and human hours have to sit on the same workstream before any of these numbers can be checked. None of them can be, today.
| Question | Evidence required | WhoWorked capability | Answerable today |
|---|---|---|---|
| Did the team adopt the practice? | Actual skill runs and outcomes by workstream, not survey recall. | Skill usage tracking | No |
| Is the practice still healthy? | Errors, drift, abandonment and unregistered variants. | Library health | No |
| Did the economics improve? | AI cost, deliverables and human hours viewed together. | Skill economics | No |
| Where did the gain land? | Project and workstream context around the assisted work. | Unified work record | No |
After four weeks of capture
None of the four is answerable today. A month of instrumented work turns each one from an estimate into a reading.
- Did the team adopt the practice?
- Runs per workstream, and which of them shipped.
- Is the practice still healthy?
- Error rate and version drift across the three skills.
- Did the economics improve?
- AI cost against human hours on the same deliverables.
- Where did the gain land?
- The share of returned capacity by work area, measured.
05 · Peer practice
What practitioners are actually doing.
Three teams further along than Northstar. Each of them found their first measure was not enough on its own.
01
Start with familiar tasks, then add context.
Moiz Imran · Senior engineering manager at Tintash · Distributed client-project teams
How they use it
The team moved from research and completion into mockup-to-code, legacy-code comprehension, database design and cleanup.
How they measure
Regular use and reported delivery and quality gains. Useful direction, not a controlled productivity result.
What Northstar can borrow
Introduce one recognisable workflow at a time and capture review effort plus the delivered outcome.
02
Usage only counts when the work improves.
Ali Dasdan and Uma Namasivayam · CTO and senior director of engineering productivity at Dropbox
How they use it
Coding, review, test generation, debugging and incident work, with internal tools where generic products lack context.
How they measure
Adoption paired with development velocity and developer-experience sentiment.
What Northstar can borrow
A frequently used tool that adds review friction still costs more than it returns. Scale only when usage and impact hold together.
05 · Peer practice, continued
The four-week experiment worth running.
One more team, then the four-week experiment that would turn this estimate into a result of your own.
03
Watch the whole delivery system.
Andrew Lau · CEO and cofounder of Jellyfish · View across 600+ organizations
How they use it
Move beyond IDE assistance into review, testing and coordination across the delivery lifecycle.
How they measure
Adoption; throughput such as cycle time and pull-request flow; then roadmap, defect and customer outcomes.
What Northstar can borrow
Faster coding with slower review, testing or client acceptance has not increased leverage.
Recommended verification
Run a four-week implementation experiment.
01
Pick one repeatable workflow.
Use the repository-change skill on comparable tickets; keep the surrounding tools and process stable.
02
Keep measures close to delivery.
Track actual skill use, elapsed cycle time, human review time, returned work and AI cost.
03
Decide with the team who did the work.
Keep what reduces total delivery effort without degrading confidence or outcomes.
06 · Notes
How the figures were produced.
The weighted task-time model behind the ratio, when the sources were last checked, and the limits of a modelled figure.
Calculation and source notes
The estimate uses the same weighted task-time model as the calculator. Each work area’s baseline share is divided by its realised multiplier, the remaining shares are summed, and the result is inverted. It is modelled from example answers, not observed team output.
Research and practitioner sources were checked in July 2026. Stable research is reviewed annually; interviews and tool guidance are reviewed quarterly or sooner.
Sources
- 01Broad field evidenceObserved · broad workplace settingpubsonline.informs.org↗
- 02Bounded implementation taskExperiment · bounded taskmicrosoft.com↗
- 03Speed versus delivered valueSelf-reported · convenience samplemetr.org↗
- 04Mature repository workObserved downside casemetr.org↗
Get this written for your firm, from your own answers.
The same eleven pages against your work mix, your team and your rate. We ask a few questions about your setup first, so the numbers are yours.