AI in Media: How to Protect Editorial Trust
In January 2025, Bloomberg began putting AI-written summaries at the top of its articles. By the end of March it had corrected at least three dozen of them. According to the New York Times, Bloomberg said that 99% of the summaries met its editorial standards. The corrections were a reminder that even when the reporting is sound, the few lines above it can tell a different story.
A reader has little reason to investigate which part of a page was written by whom. The name at the top is usually considered sufficient. Publishers have spent a great deal of money making sure it is.
By 2026, the question of whether newsrooms would use AI had largely been settled in practice. In the Reuters Institute’s 2026 trends report, 97% of the 280 news executives surveyed across 51 countries and territories rated back-end automation as important. The more difficult business is deciding where that automation should stop, who checks its work and what happens when a mistake gets through.
At Lerpal, our work on publishing infrastructure has made us wary of drawing that boundary around a tool’s name. A summariser can help a reporter navigate a long document or put an inaccurate claim above a perfectly accurate article. A tagging system can organise an archive or attach a sensitive label to the wrong story. The consequences depend on where the output goes.
So before you deploy any model, ask what happens when it is wrong. Then follow the mistake through the system. Who will see it? What decision might it influence? Can you find and correct every place it has appeared? Engineers call the reach of a failure its blast radius. It is a useful test for a publisher, provided the exercise continues beyond the first reassuring answer.
Where to Start, and What to Watch
Some publishing tasks offer a sensible place to begin because their outputs can be checked in batches and corrected without rewriting published journalism. Even here, someone needs to own the quality of the work and understand where it goes next.
Tagging, classification and metadata
An archive can contain years of valuable reporting and still be surprisingly difficult to use. Articles accumulate under old categories, new categories arrive without anyone having time to revisit the old articles, and eventually finding something depends on remembering that it exists.
Machine learning can help classify that material against a consistent taxonomy, making it easier to retrieve and match with relevant advertising. For a global digital publisher, Lerpal built a semantic vector classification pipeline that matched archived articles to the client’s advertising taxonomy. A four-person team delivered it in three months. The system combined vector search with language-model processing for cases that needed further interpretation.
That gave the publisher a classified archive it could use for more targeted advertising. It also kept the task well defined: assign categories to existing material, with no need to rewrite the articles.
The checks should follow the use of those categories. An internal filing label may be easy to correct. A label that determines advertising eligibility, appears publicly or feeds a recommendation system deserves closer attention. Test the categories that carry the greatest consequences, rather than letting a good overall accuracy score conceal the troublesome ones.
Content distribution and personalisation
Recommendation engines, personalised newsletters and “more like this” modules can help readers find reporting they would otherwise miss. They also make choices about what receives attention, which puts them closer to editorial judgment than the word “distribution” suggests.
What leads the front page is an editorial statement, and a model optimising for clicks will happily make that statement for you. If the brief rewards only clicks, the system has no particular reason to preserve the prominence of a public-interest investigation. That concern has to be part of the brief.
Editors should decide which positions they control, what coverage must remain visible and which stories are unsuitable for automated promotion. Within those limits, the model can personalise recommendations. The review should include what it repeatedly leaves out, as well as what it promotes.
Ad optimisation and revenue operations
Machine learning can inform ad density, subscription offers and paywall decisions. These systems may leave every word of an article untouched while changing the reader’s experience considerably.
A paywall model, for example, might estimate whether an article is more valuable as a subscription opportunity or as an open page earning advertising revenue. Before testing it, the publisher needs to decide which coverage should remain freely accessible and give editors a dependable way to exempt individual stories.
Public-safety coverage is an obvious case. An editor should be able to keep a flood warning open without having to persuade a revenue forecast that the flood is serious. The exemption should hold across devices, audience segments and subsequent model updates.
Search and discovery
Semantic search can help readers find an article without knowing the exact words used in its headline. Returning relevant links preserves a useful relationship: the reader can open the original reporting and see its date, sources and context.
Once search begins composing answers, the publisher has another editorial product to assess. A paragraph assembled from several accurate articles can still confuse their timelines, omit a qualification or join facts that do not belong together.
That distinction should shape the launch. A publisher can start with retrieval and links, or offer summaries that have already been reviewed. Live generated answers require a separate decision about acceptable risk, source attribution, testing and when the system should decline to answer. They cannot simply inherit the approval given to the search engine.
Transcription, translation and summarisation for internal use
Transcribing an interview or summarising a long report can spare a journalist hours of preliminary work. It can also introduce an error early enough for that error to become part of everything that follows. An internal draft has a way of acquiring authority once someone copies it into the next document.
Keep the original material close to the output. A journalist checking a quotation should be able to reach the recording and timestamp; a reporter using a summary should be able to find the relevant passage in the source. Check names, numbers and quotations before they enter published copy.
Audience attitudes leave room for these uses. In the Reuters Institute’s 2025 survey across six countries, 55% were comfortable with AI fixing spelling and grammar and 53% with AI translation. Readers, it turns out, have no sentimental attachment to typos.
There are also examples of translation supporting new publications. The BBC’s Polish-language pilot uses AI translation with a team of Polish-speaking journalists adapting and curating the material. The editorial work continues after the translation arrives.
Where AI Puts Trust At Risk
The closer an output comes to telling the audience what happened, the more demanding its review should become. Fluency makes this awkward: an incorrect sentence can be much easier to read than it is to verify.
Content generation without review
On 18 May 2025, the Chicago Sun-Times published a summer reading list of 15 books. Ten of them did not exist. The authors were real; the novels were not. The list was part of a special section supplied by King Features, a Hearst division. Its freelance writer acknowledged using AI and failing to check the results.
No Sun-Times journalist reviewed the section before it ran, and readers were never told it came from outside the newsroom. The circulation department had arranged the supplement, which appeared under the newspaper’s name. By the time anyone looked closely, the distinction between the supplier’s work and the newsroom’s work was of considerably more interest to the publisher than it could reasonably have been to a subscriber.
For anyone designing a publishing workflow, that is the detail to retain. Review requirements need to cover every route to publication, including licensed supplements and commercial arrangements made outside the newsroom. A careful process on the main editorial desk cannot protect a page it never receives.
Reader-facing summaries belong here too. Bloomberg’s summaries sat on top of accurate, human-written articles and still needed correcting, because a summary is new text with the outlet’s name on it. It needs its own accuracy check.
Disclosure and responsibility
In the Reuters Institute’s 2025 six-country survey, comfort with news production rose with human involvement: 12% were comfortable with entirely AI-generated news, 21% with AI-led news under some human oversight, 43% with human-led news assisted by AI and 62% with entirely human-made news.
Those findings support being specific about who did what. “Translated with AI and checked by an editor” gives a reader useful information. “AI-assisted” leaves them to guess whether the assistance involved a spelling correction or most of the article.
A disclosure should describe a process the publisher can account for. If it says an editor checked the work, there should be a record of that review. A label cannot carry responsibility on behalf of a person who was never given time to exercise it.
Automated editing without guardrails
AI editing tools that rephrase, restructure or “improve” copy can change meaning without producing anything that looks obviously wrong. “May have caused” becomes “caused”; an allegation loses the words that identify it as one. The sentence is shorter, smoother and less accurate.
Summarisation deserves the same caution. In the BBC/EBU’s 2025 study of consumer AI assistants, 45% of the more than 3,000 answers assessed had at least one significant issue. These were third-party assistants, rather than publishers’ internal editing tools, so the figure should not be treated as a newsroom error rate. It does show why plausible wording and links to sources are insufficient evidence of an accurate answer.
For published copy, editors need to see the changes alongside the original. Names, figures, quotations and qualifications deserve particular attention. A headline should be checked against the reporting it introduces, however well it performs in a test of alternative wording.
Build the Review Into the Publishing System
An AI policy is often written in calmer circumstances than those in which it will be used. At 11:40 pm, with a story still developing and several versions of the headline in circulation, the publishing system needs to make the next responsible action easy to find.
For CMS teams, we would start with four requirements:
- Record where content came from. Each paragraph, headline and summary should retain its origin and revision history, including whether AI or an outside supplier contributed. Preserve that information when content moves between systems.
- Require approval for the version that will go live. Anything drafted or substantively changed by AI should need a named editor’s approval. If a model rewrites it after approval, that approval no longer covers the new text. Give third-party editorial material an explicit review route too.
- Keep an audit log. Six months later you should still be able to establish what was published, who approved it and which model and configuration produced the relevant output.
- Make each feature easy to stop. When a summary tool starts misbehaving, an authorised person should be able to disable it without a software release. Decide what readers will see instead, and how the team will locate outputs already published.
These controls need to match the product. An editor can approve a stored article summary before publication. They cannot approve an answer that will be generated next Tuesday in response to a question nobody has yet asked. If pre-publication review is a requirement, the feature must work within it.
The reviewer also needs a workable job. Show the source and the proposed output together. Make substantive changes visible. Allow enough time for verification, especially when a fluent paragraph conceals several claims.
Review times and edit rates can help identify where to look more closely, but they require interpretation. An unchanged draft may be accurate; a heavily edited draft may still contain the original factual error. Periodically inspect the quality of the review itself, with the editors who perform it.
Give Each Feature a Measurable Job
Start with a task whose output you can evaluate and whose mistakes you can contain. For tagging, that might mean testing a defined part of the taxonomy before processing the full archive. For search, it might mean checking whether readers find the right reporting. For transcription, count the time spent correcting the transcript as part of the cost.
The Reuters Institute’s 2026 findings warrant some restraint about returns. While 64% of the executives surveyed rated back-end automation very important, only 13% described their newsroom AI initiatives as transformational. Forty-four per cent called the results promising and 42% limited. Importance and demonstrated value are different measurements.
Choose the measure before the pilot starts. Time saved should include review and correction. Revenue tests should also track the effects on reader experience. A feature that produces twice as much material has created twice as much material; whether it has improved the operation remains to be established.
Monitor AI outputs continuously
A successful pilot tells you how a system performed under the conditions of that pilot. Models, prompts, source material and newsroom practices can all change afterwards.
Keep a fixed set of examples to test for regressions, and add fresh samples from production to catch problems the original tests did not anticipate. Review performance when a model or configuration changes, with particular attention to names, dates and developing stories.
Set thresholds and responses before launch. A tagging error might call for correcting a batch and reviewing one category. A fabricated quotation should trigger an immediate investigation into the affected workflow. Decide who can pause the feature, who corrects published material and how those corrections reach copies distributed elsewhere. An overall error percentage should never conceal the severity of a particular failure.
Train editorial teams on AI literacy
Journalists don’t need to become engineers. They need to understand where the tools can mislead them and how to check the work efficiently. Training should use the newsroom’s own material: an interview with a difficult name, a report with conflicting figures, a story whose meaning depends on one carefully placed qualification.
Include the commercial and circulation teams. At the Sun-Times, the section that did the damage was bought outside the newsroom. Anyone who can commission or introduce material for publication needs to know the review requirements.
There should also be room to say that a tool is creating more work than it saves. A reporter who spends longer checking a summary than reading the source has learned something useful about the product. The pilot should be allowed to learn it too.
What the Reader Should Be Able to Expect
Readers are unlikely to follow the details of a publisher’s model choices or approval system. They should be able to expect accurate reporting, a clear account of substantial AI involvement and corrections when something goes wrong. The publisher remains responsible for delivering those things, including when the mistake began with a supplier or a seemingly minor automated edit.
That responsibility has practical consequences for the software. It determines what can publish automatically, what needs review, what information the reviewer receives and how a feature can be stopped. These decisions belong in the design from the beginning, while changing them is still relatively inexpensive.
At Lerpal, we work on the systems that support those decisions. Our collaboration with Hearst Digital Magazinesdates to 2020 and includes CMS development, search and publishing infrastructure. That experience informs how we approach AI projects: define the task, establish who is responsible for its output and make sure the workflow can put a mistake right.
If you are weighing an AI feature for your publication, bring us the workflow you’re considering. We can help work through what to automate, where review belongs and what the team needs when the system gets something wrong.
You may also like
Talk With Our Experts
Get advice and find the best solution