Structured AI Development · sourced snapshot

By the
Numbers.

A production story, not a scoreboard.

In one month, a human-led AI organisation moved from initial concept into active production, validation and publishing preparation. The useful question is not which number is biggest. It is how the numbers connect.

Start with authority ↓
Human Supervision Required Studios editorial evidence mark
1 MONTHfrom initial concept to this sourced baseline [4]

01 / Control

Scale starts with a boundary.

Before counting AI roles or outputs, the project defines who can actually decide. The answer remains deliberately singular.

Final authority 1 HUMAN

holding final product, creative, spending and release authority. [12]

14human-defined non-negotiable product directionsSource [12]
8standing operating rules controlling how the AI organisation worksSource [12]

02 / Organisation

One authority. Many jobs.

Capacity comes from dividing the work into explicit disciplines rather than asking one general-purpose assistant to impersonate a studio.

Final authorityHUMAN
DesignEngineeringProductionArtAudioQAPublishingBrandFinanceExecution
10named specialist AI roles across studio, project and execution disciplinesSource [13]
17development functions tracked from foundation through post-launchSource [10]
2 operating layersstudio governance + project execution[13][12]

03 / Output

The work has to leave a trace.

Chat volume is not the output. The useful count is durable work that survives into governed project records, production systems or approved content.

Durable production record351

durable project outputs recorded by the operational ledger. [8]

Accepted outcomes148

approved outcomes recorded by the operational ledger. [8]

638approved authored runtime dialogue lines in the current content baseline[6]
540validated response-state combinations in one interaction system[7]

04 / Test

Then test the claims.

Generation speed is not confidence. Repeatable validation, deterministic checks and coverage tests are the evidence that a system behaves as intended inside its stated boundary.

Controlled synthetic testing31,000validation runs across recorded test programmes [1][2]
4,368validated scenario configurations in the replayability modelSource [3]
230 / 230deterministic behaviour-verification cases passedSource [9]
40 / 40required behaviour-coverage cells executableSource [9]

05 / Comparator

Big numbers need big caveats.

The project also reconstructs a conventional-development comparison model. These figures are useful for planning and hypothesis testing. They are not recorded human hours, booked cash savings or proof of final project economics.

Reconstructed scope comparator3,068conventional-development hours represented by the governed baselineComparator · not recorded human timeSource [4]
Schedule comparator87calendar-day lead against the standard-team comparatorModelled comparison · not a release claimSource [4]
Replacement-cost comparator£112,667base reconstructed replacement-cost comparator at the last closed checkpointNot claimed cash savingsSource [5]

The number we won't invent

Proof of success.

These records demonstrate production activity, structure and bounded technical evidence. They do not yet establish market demand, external player response, final release quality, total production cost or that the same results will repeat unchanged on another project.

That is the next evidence problem.

Human Supervision Required Studios — Machines assist. Humans decide.

References

Every claim keeps its source.

Reference numbers link the public-facing figures to controlled project evidence. Public labels are deliberately sanitised; supporting material can be shared at an appropriate level for a serious enquiry.

Human Supervision Required institutional evidence stamp Request supporting evidence
  1. Synthetic validation record. A 30,000-run controlled validation set retained as valid within its stated boundary.
  2. Diagnostic simulation record. A separate 1,000-run controlled diagnostic set used for repeatable system evaluation.
  3. Replayability validation record. 4,368 valid scenario configurations recorded under the approved model rules.
  4. Production timeline baseline. Records the project start point, the 3,068-hour reconstructed scope comparator and the 87-calendar-day schedule comparator.
  5. Production cost comparator. Last closed checkpoint: £112,667 base reconstructed replacement-cost comparator; not a claim of cash savings.
  6. Approved content baseline. 638 authored runtime dialogue lines; pre-generated rather than created live at runtime.
  7. Interaction-system validation baseline. 540 linked response-state combinations in the reviewed system.
  8. Operational output ledger. Through the snapshot date: 351 durable outputs and 148 approved outcomes.
  9. Behaviour coverage verification. 40/40 required coverage cells executable; 230/230 deterministic verification cases passed.
  10. Development lifecycle baseline. 44 lifecycle objectives across 17 tracked development functions.
  11. Evidence index. At least 32 key governed evidence sources indexed in the reviewed baseline.
  12. Product authority record. Fourteen current core product directions, eight standing operating rules and explicit human final authority.
  13. Organisation and role-governance records. Ten named specialist AI roles across studio, operational and task-bound execution.
  14. Core project-control records. Four independent control records covering progress and cost, status and issues, studio chronology, and system/service outcomes.