Regulators want Columbo, not Sherlock Holmes: the conclusion first, stated plainly, followed by exactly enough evidence to trust it - not every data point in the order it happened.
Merck's collaboration with McKinsey reduced the time to produce a fully human-reviewed first-draft CSR from an average of 180 hours to 80 hours, while reducing reported errors by 50%. The overall time to produce the CSR fell from two to three weeks to three to four days. (Source: Merck) That result makes one thing increasingly clear: AI can materially accelerate regulatory document authoring.
Which is exactly why the more interesting question isn't being asked in speed terms at all anymore. It's the one Dr. Barry Drees has been putting to medical writers in workshops for years, long before "AI governance" entered the vocabulary: give us Columbo, not Sherlock Holmes.
It's a reference to two kinds of detective story. In a Sherlock Holmes plot, the reader only understands what happened on the final page. In a Columbo episode, you see the murder in the first five minutes - the tension is in watching the case get built, not the reveal.
Barry's point is that regulators want Columbo: the conclusion first, stated plainly, followed by exactly enough evidence to trust it - not every data point in the order it happened.
Ask most medical writers why documents run long, and the honest answer isn't laziness. It's fear. Leave something out, and it might come back as a Request for Information - so the instinct is to include everything: every number restated in a table and again in prose, every possible confounder addressed at length whether or not it matters.
Barry calls this "fat writing" - and simply cutting word count without cutting substance isn't much better. The real target, in his words, is "Story, Not Storage." Not a warehouse of every fact that could conceivably be relevant. A document that tells the reviewer what matters, and why, and then gets out of the way.
The reason lean writing is hard to sell internally is that most writers have never seen how their documents are actually reviewed. A submission doesn't land on the desk of the world's foremost expert in the therapeutic area - it's read first by reviewers with a checklist, verifying every number against source data, before any specialist committee sees it distilled.
That changes the calculus completely. A thousand pages of over-qualified text doesn't protect the sponsor. It creates work the reviewer can't resolve, increases the odds the important sentence gets lost, and, often, leads reviewers to simply stop reading in full. In a 2026 review published in DIA's own Therapeutic Innovation & Regulatory Science, researchers at Eli Lilly documented direct feedback from FDA reviewers: overly long clinical documents were sometimes skipped over entirely, occasionally contributing to delayed or missed approval opportunities. (Source: Ther Innov Regul Sci, 2026) Length doesn't read as thoroughness to a reviewer working through a checklist - it reads as a document they can't fully get through in the time they have.
Part of the problem is that "the guidance" isn't one thing. The same paper describes mapping ICH E3 against the wider set of guidelines that govern efficacy and safety reporting, and against the eCTD structure those documents ultimately populate - so a writer knows exactly which document a given piece of content belongs in before drafting starts. Without that map, a writer facing an ambiguous, cross-referenced set of guidelines defaults to including everything, just in case.
This is where the conversation about AI in regulatory writing usually goes wrong. The industry has largely settled the question of whether AI can generate clinical text - it clearly can. The much harder, much more consequential question is: what should it write?
An AI model with no defined standard for "good" will default to whatever's most common across the examples it's seen - all the fat writing, all the repetition, all the caveats added out of fear. Without a clear target, AI doesn't fix the over-writing problem. It scales it.
That's why standardisation has to come before automation, not after. Before a model can reliably write the sentence "there were no differences between treatment groups that would affect the interpretation of safety and efficacy," something upstream has to have already determined, with evidence, whether that's true. Only once that judgement is established should any language be generated at all.
This isn't a debate medical writers are having only with themselves. DIA's Medical Writing & Scientific Communication Conference this year built a session around a "protocol-first" approach enabled by ICH M11 (Source: DIA) - using structured, machine-readable protocols as a single source of truth. That's standardisation-before-automation, stated as a conference agenda item. The DIA Global Annual Meeting ran a parallel session on the "AI Competency Framework for Medical Writers" (Source: DIA), framing the discipline's future around writers becoming AI stewards - people responsible for defining what AI should and shouldn't be trusted to decide.
Regulators are converging from their side, too. In January 2026, FDA and EMA jointly published ten Guiding Principles of Good AI Practice in Drug Development (Source: FDA), built around traceability, accountability and risk-proportionate validation.
Push that logic forward and one scenario gets uncomfortably plausible fast: sponsor AI generating language that agency AI then reviews, with no shared, explicit standard sitting between the two. Two probabilistic systems checking each other's work is not oversight - it's two guesses compared against each other, with no fixed point either one is answerable to. The only way to avoid it is agreeing on the standard - the finding, the sentence, the structure - before any AI is allowed to write anything at all.
The industry has already published a playbook for this. In June 2026, a team at Eli Lilly, together with medical writing leaders from Novo Nordisk, Sanofi, Roche, GSK, Pfizer, Novartis and Merck, published a detailed framework in the same journal. (Source: Ther Innov Regul Sci, 2026) Their central claim: lean authoring isn't a style preference, it's infrastructure - and it has to be built before automation, not layered on after. The results are concrete: Merck saw roughly a 50% reduction in CSR length; Novo Nordisk cut the time from data availability to CSR approval by as much as 90%.
None of this is entirely new. A decade earlier, medical writers hit the same problem with no AI involved: CSRs were growing longer, more repetitive, harder to review. The response was CORE Reference (Source: EMWA) - a pro bono AMWA/EMWA standard for writing CSRs that communicate what matters without burying it. It's still in active use, and one of the sources the 2026 framework builds directly on. What's different now is the stakes: getting the standard right is no longer just about better human writing. It's the prerequisite for whether AI amplifies good practice or bad habits at scale.
This is the exact principle TriloDocs is built around. Our deterministic reasoning engine identifies clinically relevant findings first - cross-checking AE frequency against severity, ruling out demographic confounding, reconciling tables against each other - before any generative model is involved. Only once a finding is established does an LLM render it into language.
The model doesn't decide the science. It expresses what's already been decided. Every sentence in the output traces back to a specific rule, a specific piece of evidence, and a specific reasoning step a reviewer can inspect.
AI can make medical writing faster. But speed is only useful if the process starts with knowing what needs to be said, why it needs to be said, and what evidence supports it. That's the principle TriloDocs is built around: story first, language second.