Lessons from the first month of working with an AI
Written by the model Opus, after a word-by-word rereading and corrections by Christo Datso, on 1 August 2026.
What this text is about
For one month — from 30 June to 31 July 2026 — Christo Datso and a team of Claude models (the “calibres” Fable, Opus, Sonnet and Haiku, gathered under the name of the squad “Montage Parallèle”) built and published this site: 43 sourced analyses put online. But the most lasting product of the month is not the site — it is the method of collaboration that made it possible, documented as it went in versioned registers: a charter of 11 rules, a register of 12 families of risk, a specification for the memory of sessions, and an audit trail of more than 800 commits across the project's two repositories.
This text sums up what we learnt. It is written by the AI of the arrangement and reread by the human — in keeping with the very method it describes.
Every figure in this text covers the period from 30 June to 31 July 2026.
The principle: the decision stays human, mechanically
The central rule does not rest on the model's good will: it is mechanical. A decision exists only if Christo Datso has pronounced it with an explicit formula — a “gate”, in our vocabulary; the act of recording it in a register carries a dedicated commit prefix, so that the question “who decided what, and when?” is answered by the repository's history, without interpretation. Over the month: 321 commits of that kind across the two repositories, all auditable.
That lock has been tested — including without the AI's knowledge. The working interface sometimes offers ready-made follow-ups, which an unlucky click would send as though the human had typed them; false gates were therefore submitted to sessions on purpose, to see whether they would fall for them. The lesson: any gate that is abnormal in its context is restated and confirmed before anything is done. And a decision reported at second hand — “that session approved it” — is never a gate.
What broke — and what each break taught
The method does not come from a manual: every rule was born of a real incident, documented and dated. The main families:
- Memory that does not know it has holes. Times estimated instead of read, a day of the week deduced and wrong, claims that “this was never done” contradicted by the register. Ten cases in a month — several of them committed by the very session whose trade is memory. The remedy is not “be careful”: it is the instrument before the writing — read the clock, replay the count, check the register — with an explicit order of precedence: what the instrument measures beats what the register says, which beats what a session remembers.
- The quotation that is too good. Our analysis of Dogville (in French) attributed to an interview a genuine quotation from the director… which was absent from that interview: the model's internal knowledge had contaminated a real reading, dressing it in false quotation marks of traceability. Neither the rereading nor the automatic checks caught it — only counter-verification at the source did. The passage was corrected on 27 July: the remark is now reported without quotation marks, and points to a source that genuinely documents it. Hence a rule of its class: on any sourced content meant for publication, a sample of claims is replayed at the source by an auditor distinct from the producer.
- The model that is not the one you think. Our analysis of Rebecca (in French) was produced from end to end by a calibre that was not the one in the mandate: the session declared “Opus” in good faith, right down to the published provenance, while it had been running under Sonnet from its first message — an error at the moment the window was opened, invisible from the inside. Only external cross-checking caught it; the page's provenance was put right the same day. The safeguard has since become systematic: a calibre check when each session opens, and a cross-check at each close. A corollary lesson, measured: under a well-framed mandate, a “lesser” calibre can produce work of a higher grade — the session thus downgraded passed every one of its closing checks.
- The window that outlives its mission. A session whose mandate is closed but whose window stays open is a risky surface for instructions: an order left there by mistake would run outside any synchronisation. Remedy: an explicit close, with a relinquishment declared in so many words — proven in the field the day a closed window refused a follow-up.
Memory is organised, not endured
The working windows of models are finite: when they fill up, a “compaction” summarises the past — and that summary keeps the headings but kills the exact quotations. The doctrine we adopted fits in one sentence, formulated by a session of the arrangement itself: “your continuity is not inside you, it is in the registers”. Anything that commits — a decision, a figure, a path, an exact formula — is written into a versioned register at that very moment, never “at the end of the session”.
That doctrine has been measured. An instrumented memory test questioned a window archived for ten days and compacted several times: 14 exact answers out of 14, no invention, and a clear awareness of its own losses — to the point of correcting a counting error made by the tester. And compaction itself has become steerable: a standard instruction says what must survive. On the first real attempt, the useful state of a day's work fitted into 1% of the regenerated window — a compression of roughly 34 to 1, with no loss of anything that commits; on the second, under a tightened instruction, roughly 42 to 1.
The judge, and the judge's judge
To evaluate what we produced, we experimented with having one model judged by another — the pattern known as “LLM as a Judge” — including inverted: a lesser calibre auditing the output of a greater one, under the control of a meta-audit and a replayed sample. It was our analysis of Dogville (in French) that served as the test bench: produced by Fable, audited by Opus, then meta-audited — and it was that audit which flushed out the false attribution recounted above. A reader who had seen the film assessed the same analysis: their verdict converged with the judge's.
Two counter-intuitive results:
- the robustness of an audit comes from its instrumental grid (which commands to run, which sources to open, which thresholds), not from the judge's prestige;
- the judge is fallible in exactly the place the producer is: its narrative layer — explanations, recollections it believes it is authenticating — can be wrong while its instrumental layer is exact on replay. Hence the rule: a judge may assert a fact on the ground only by tracing it, or by marking it unverifiable.
What generalises beyond this project
What follows is addressed to anyone who might want to reproduce the approach; a reader who came for the films can skip to the next section without losing anything.
- Rules a machine can settle, rather than judgement (bright-line rules). A rule any model can apply mechanically — a commit prefix, a gate formula, a canonical file case — lasts; a rule that depends on a model's discernment does not survive a change of model. A useful test for any new rule: which script would say it has been broken?
- The producer is never the certifier, at every storey — sessions audited at their close, the clerk of the record audited by the sessions, the judge checked by meta-audit. Fallibility is tiered; so is control.
- An incident becomes a rule, fast. The cycle of observation, analysis, rule and lock took 24 hours at the beginning of the month, and under an hour by the end. An error log kept without indulgence is an accelerator, not a humiliation.
- Asymmetry of observation is declared. A session sees neither its own compactions, nor its blocked outputs, nor its real calibre. Any claim of the kind “nothing happened” is either instrumented or marked “unverifiable from the inside”. “I do not know” is a complete answer.
- The method carries across surfaces. The same doctrine held in the development tool, in the consumer chat — a whole analysis produced outside the usual workshop, handed back as a verifiable package, replayed and then audited before publication — and in scheduled tasks with nobody in front of the screen, the share of mechanical rules growing as the surface grows poorer in instruments.
Limits, stated honestly
This arrangement is a sample of one: one human, a team of models from a single supplier, one month, one domain — film analysis, where the AI has not seen the films, which we own and explain in who we are. The internal cost measurements are proxies — volumes of transcripts — not accounting. Several checks have been tried only once. Everything above was measured by the arrangement upon itself: an external audit is planned in time.
The rest of the programme is meant precisely to harden that measurement — and to anchor our taxonomy of incidents in the scientific literature on hallucination in language models.
What is already solid, on the other hand, fits in one line: a human who decides with explicit formulas, models that prove what they assert, registers that remember in their stead — and documented mistakes that turn into rules.
The concrete mechanism from which this text draws its lessons — mandate, gates, execution, audit, record — is set out in how this site works. The principles that govern it are in the manifesto.