Building a better news engine: how we came up with Stein
In May we proposed seven article types you could read off a URL. Then we tested it. This is what broke, what we rebuilt and what Stein is.
Every taxonomy is elegant until it meets a dataset.
Last May, we proposed a seven-type framework for classifying news articles built on a simple yet complex premise: all that an LLM needs to categorize an article is contained in just three elements. This statement is supposedly true for the whole internet, where URLs, titles, and images are full of decipherable signals and rich, interpretable information.
The limit of a retrieval system that takes just those three elements into account is that while they are all part of the same output—an article—they are likely conceived and generated by different people with different tasks and goals. Newsrooms are big, complicated, and industrialized functions of the media business: e.g., the person who writes the article is not the one who creates the title, while someone else entirely might work on the lede, the images, the internal text, etc.
The media business is just... a business
One thing I’ve learned while working in the media industry is that titles and content serve two fundamentally different purposes: titles are for conversion, while the rest of the content is for engagement. Publishers are in the attention business and play by the same rules as everyone else: each piece of content has to be “thumb-stopping” and interesting. Once you start reading it, the body has to really hook you in. Images must reflect the article content, but they must also be evocative and cool—something you are willing to share on your social media profile. Internal links must be useful to keep the reader clicking through to different things until they eventually want more and subscribe (it’s much more complicated than this, but that’s the point).
For this very reason, every article is just a complex atom of an industrial machine where everything is taken into consideration. The media plays a pivotal role in society, but the media business is mostly just... a business. Not dissimilar to how retail works: A store needs a great window to lure people in, the products must be organized in a certain way to keep you inside, and the packaging and salespeople make you convert. Just like an article does.
Introducing stein
Our goal as a research group is to fight polarization by providing an open-source technology built in the most transparent way. The first step to doing so is to categorize articles and media outputs into clear categories, which is what we tried to do in May with a first iteration of our framework. But as soon as we got access to a more complete dataset, our hypothesis was confirmed: the core elements of an article don’t always match. A title is a premise, the lede locks you in, and the content of an article could be something completely different.
This is how we built the second iteration of our open-source framework, which we named Stein—after Gertrude Stein, the great writer, poet, and mentor to many iconic authors (like Hemingway) of the last century (you can learn more about her here).
The framework (including the taxonomy, the classifier prompt, and the validation methodology) is available in our repository [https://github.com/UnbubbleHub/Stein].
We couldn’t have done this without the help of our friends at Column, the AI-powered news platform that helps readers cut through the noise by bringing together high-quality journalism from hundreds of trusted publishers into a personalized, distraction-free reading experience. Their team generously provided the clean, comprehensive dataset that made this evaluation possible and collaborated closely with us throughout the project. Since improving how news is organized and discovered is at the core of what they build, they were the ideal partner for validating our framework against a real-world media ecosystem.
More on Unbubble Hub
Unbubble Hub is an Open Research Initiative that provides a space for researchers and engineers to come together and collaborate in developing tools to fight social polarization.
Sources is a GitHub repository (a piece of code) that takes a news event and returns sources, categorized and ranked, representing a range of diverse viewpoints.
Giorgio Catalani (find him on LinkedIn) started his career as a journalist and slowly moved closer to technology. He works as a sr product manager for a major publishing group; everyday, he tries to make life easier for editors and journalists. Big internet nerd, movie buff and wannabe cook. This is his first article for Unbubble Hub.
Signals are promises, not truths
As noted, we ran our first experiment with a simpler dataset, but as soon as we got access to a complete and diverse dataset, we started hand-labeling a few hundred articles. We immediately noticed what we already suspected to be true: titles and URLs are not enough.
A signed column could run under a hard-news headline in the sports section, a product launch gets written up as a leisurely walk-through, an Italian daily puts a full-throated editorial argument in Economia because that’s where the desk sits, not where the genre lives. Esser and Umbricht saw this coming a decade ago: their content analysis of 2,422 political stories across six Western press systems found that even “pure news items” have been steadily colonized by interpretation since the 1960s—the genres bleed into each other inside the text, invisibly to any metadata.¹ You cannot referee a boundary dispute you cannot see.
This confirmed our hypothesis that titles and URLs are not enough: the body of an article is where the ground truth lives, while titles and URLs are demoted to priors and tie-breakers. The classifier we have now reads the full text and rules on what the piece actually is, structurally; when the headline and the body disagree, the body wins.
Roughly a fifth of our hardest validation cases were exactly this: an article whose title, read alone, would have sent it to the wrong bucket.
This sounds obvious (and it sort of is), but it is also expensive and complex if you work with RSS or simple RAG systems, which tend to fetch just titles, URLs, and a small snippet of text. We’ll get back to this at the end when we talk about our challenging next steps.
From seven types to three postures
The second casualty of the real-world dataset was the flat list itself. In the first iteration, we had seven simple categories standing side by side. This implied that confusing an Interview with an Opinion was the same kind of mistake as confusing Breaking News with a Liveblog. It isn’t. One error changes the epistemic nature of what the reader gets.
So, the taxonomy is now more complex and split across two levels. At the top are three postures representing the article’s stance toward reality:
FACT: The piece reports who, what, when, and where. It is attributed, with no interpretation in the writer’s own voice. (Types: Breaking News, Quote, Liveblog)
FRAME: The piece interprets. It explains, contextualizes, quantifies, or verifies empirically—without prescribing. (Types: News Analysis, Explainer, Poll Report)
VIEW: The piece takes a position. It offers a normative judgment or recommendation in the author’s own voice. (Types: Opinion, Review)
We “demoted” some of the old framework categories. Wire Republication turned out to be a matter of provenance, not genre: the same Reuters dispatch is Breaking News whether it runs in one outlet or six. Fact-Check turned out to be a structural variation of News Analysis (claim → evidence → verdict is interpretation with a scorecard). Interview collapsed into a broader type, Quote, once we noticed the real distinction wasn’t the Q&A typography but something deeper: whether the declaration itself is the news.
The macro/type split isn’t cosmetic. It encodes a judgment about which mistakes matter. If the system files a Liveblog as Breaking News, the reader still gets facts. If it files an Opinion as News Analysis, it has laundered advocacy into interpretation—and that is precisely the failure a diversity engine cannot afford.
For context, here is the full second level of the framework:
FACT — The piece reports
Breaking News: A new event or development is reported in the outlet’s own voice: a launch, ruling, appointment, result, death, or deal. Quotes may fill the piece, but strip them out and something still happened.
Quote: The declaration is the news. Nothing occurred beyond the saying of it: an interview, a communiqué, a minister’s warning, or an analyst’s forecast. Remove the quotes and nothing remains but “X said Y.”
Liveblog: An unfolding event covered via reverse-chronological, time-stamped running updates. The format is the point: short rolling posts, newest first.
FRAME — The piece interprets
News Analysis: The writer advances an interpretive thesis about an event the reader is presumed to already know: why it happened, how, and what it means. It is empirical, not normative. This includes claim-checks and debunks (claim → evidence → verdict).
Explainer: Pedagogical and utility-oriented: what is X, how does it work, and what changes for you. Backgrounders, biographical profiles, timelines, walk-throughs, and curtain-raiser previews all live here. It is neutral by construction; it teaches, it doesn’t argue.
Poll Report: Reports survey results as data: percentages, sample size, margin of error, and crosstabs. The interpretation is carried by the numbers, not by a thesis.
VIEW — The piece takes a position
Opinion: An argument in the author’s own voice: an editorial, a column, or a position on an issue. The object of the piece, if there is one, is a springboard to say something about the world beyond it.
Review: An evaluative verdict bound to a specific work, product, or event. It judges the thing’s own qualities, usually with a rating or a consume/skip recommendation. Even a partial slice (“6 features I like”) counts, as long as the judgment stays focused on the object.
P.S. Sports content is becoming really hard to research, as most of it focuses on odds, predictions, and betting. We’ll eventually circle back to that in a later piece.
The tests that fell out of the labeling
The genuinely useful outputs of a labeling exercise are never the labels themselves. They’re the tests you’re forced to invent when two humans disagree. Three earned their place in the classifier prompt:
Whose voice interprets? A news report that quotes an economist’s analysis is still FACT—it reports that the economist analyzed something. It becomes FRAME only when the publication’s own writer interprets the situation in an analytical voice. This single distinction resolves most FACT/FRAME disputes, closely matching Salgado and Strömbäck’s operational view: interpretive journalism is defined by the journalist’s explanations, evaluations, and speculations going beyond verifiable facts, not the sources’.²
The remove-the-quotes test. Take a piece and strip every direct quote from it. If a reportable event still stands, it’s Breaking News. If nothing remains but “X said Y,” it’s a Quote. A minister’s demand or warning is a verbal act, not an event.
The lede test. If the first paragraph reports the event itself, the piece is Breaking News, even if paragraphs four through nine interpret it. If the piece presupposes the event and immediately pivots to explaining it, it’s FRAME. The lede is where the author declares their primary intent, whether they mean to or not.
The numbers, methodology, and models
We explored a few hundred hand-labeled articles, froze a 120-article gold set for validation, and ran the classifier across roughly 1,700 articles in production. The set spans four languages—English (52), Italian (46), German (12), and Spanish (10)—and covers 18 deliberately different, high-traffic topics, including politics, tech, finance, culture, and sports.
A first batch of 49 was gold-labeled entirely by hand, while a second batch of 71 was adjudicated with Opus 4.8 on high effort to propose labels, with a human ruling on each. Ultimately, the final gold set overrides Opus’s own call in four of those cases. This is the difference between true adjudication and rubber-stamping, but it still means most of our gold set had a model in the room. Keep that in mind for what follows.
The production classifier, running on full text, agrees with the gold set on the macro posture about 89% of the time, and on the exact type among the eight categories about 83% of the time. This performance was nearly identical across both the hand-labeled and adjudicated batches. Opus itself, on the uncontaminated hand-labeled batch, reaches about 96% on macro postures and 88% on specific types. Regarding that last number: the strongest model we had beats our production classifier by only four points on fine-grained distinctions. These types are hard for everyone. On the full set, Opus scores higher, but because it helped adjudicate that batch, it is partly grading its own output; therefore, we do not report that number as a definitive ceiling.
Two patterns in the errors are worth reporting:
First, the errors are worse than they look, but better than we feared. Of twenty type-level disagreements, thirteen crossed postures, mostly along the FACT/FRAME seam (e.g., Breaking News filed as an Explainer, or a Quote filed as News Analysis). This is the boundary line between reporting, interpreting, and teaching around the same event. Two cases are textbook: a signed Italian piece using the World Cup to argue about Trump’s America, and a German morning briefing arguing the case against a stock. While the classifier sometimes confuses different kinds of analytical honesty, it almost never mistakes outright advocacy for objective evidence.
Second, we thought that the FACT/FRAME seam would be hardest in interpretive press cultures, especially in Italian. It isn’t—or at least, isn’t yet. English and Italian error rates are statistically indistinguishable at this sample size, and our German sample is too small to say anything definitive. Either the classifier’s structural tests genuinely neutralize press-culture differences, or 120 articles aren’t enough to expose them. We’re keeping the question open, which brings us to the literature.
The framework was hiding in sixty years of journalism studies
When we proposed the seven types in May, we cited the comparative-journalism canon in footnotes. Having now labeled a corpus, I want to promote those footnotes to the main text because the correspondence is closer than we realized.
Esser and Umbricht’s six-country study coded political stories into four categories: news items, information mixed with interpretation, information mixed with opinion, and commentary.¹ Squint slightly and you have FACT, FRAME, and VIEW (with their two middle categories marking the contested borderland our classifier’s confidence scores now try to measure). We arrived at the same shape from the engineering side by asking which disagreements between human labelers refused to die. The convergence is reassuring: it suggests the postures are real properties of texts, not artifacts of our prompt.
Salgado and Strömbäck’s review of interpretive journalism is, in retrospect, the specification document for our FRAME macro.² Their central complaint was that the field had no consistent operationalization of “interpretation,” making studies incomparable. This is exactly the type of problem that an LLM classifier forces you to solve because a model cannot run on an ambiguous definition.
Reinemann, Stanyer, Scherr, and Legnante did the same disciplinary work for hard/soft news, concluding that the distinction only becomes usable when decomposed into separate dimensions of topic, focus, and style.³ We took that lesson to heart: our taxonomy deliberately classifies postures while remaining agnostic about topics. A soft-topic piece (about a restaurant or a TV series) can be pure FACT, and a hard-topic piece on monetary policy can be pure VIEW.
And Hallin and Mancini’s polarized pluralist model⁴ explains our per-language results before we even ran them. Comparative studies report the share of interpretive stories ranging from around 31% in the UK to around 60% in Italy.⁵ Italian journalism blends reportage and commentary as a matter of tradition, not sloppiness. This is why we insisted from the start that the Italian ecosystem would be our stress test, and why the FACT/FRAME seam is where our classifier earns or loses its keep.
There is one more shelf of literature that didn’t exist when Esser and Umbricht were coding by hand: recent work validating large language models as content-analysis annotators. Gilardi, Alizadeh, and Kubli found LLM classifications more consistent with trained coders than crowd workers across several annotation tasks.⁶ Törnberg reported similar results against both experts and crowd workers on political text.⁷ We are, in effect, running a small, applied replication of that research program.
What’s next
Here is what we are focusing on in the coming weeks:
Keep studying the framework: One hundred and twenty articles is a validation set, not a final verdict (our German sample, in particular, is too small to analyze confidently). We want a larger corpus, more languages, better calibration of the confidence scores and a proper treatment of genuinely mixed-mode pieces.
Make it work with just the title and the text: Today, the classifier leans on the URL as a prior, but URLs are the least portable part of the input. They vanish in RSS payloads and API responses, and they get mangled by syndication. Our goal is to build an open-source technology that can be used as easily as possible. If the framework is going to be useful beyond our own pipeline, it has to rule on any article anyone hands it. We will try to make it work with the fewest elements possible.
Trace how topics move through time: Our underlying thesis is that a topic moves through time the way a product does. It starts as an MVP (a breaking news item) and, if it gains traction, it travels across the framework as readers engage, ask for context, and finally seek judgment. FACT begets FRAME begets VIEW. We want to measure that lifecycle (again, because the media business is just... a business).
Open source: The framework—including the taxonomy, the classifier prompt, and the validation methodology—is available in our repository: [https://github.com/UnbubbleHub/Stein].
As always, we’ll publish what we learn, including the parts that break.
References
¹ Frank Esser & Andrea Umbricht, “The Evolution of Objective and Interpretative Journalism in the Western Press: Comparing Six News Systems since the 1960s,” Journalism & Mass Communication Quarterly, 91(2), 2014. ² Susana Salgado & Jesper Strömbäck, “Interpretive Journalism: A Review of Concepts, Operationalizations and Key Findings,” Journalism, 13(2), 2012. See also Salgado, Strömbäck, Aalberg & Esser, “Interpretive Journalism,” in Comparing Political Journalism (Routledge, 2017). ³ Carsten Reinemann, James Stanyer, Sebastian Scherr & Guido Legnante, “Hard and Soft News: A Review of Concepts, Operationalizations and Key Findings,” Journalism, 13(2), 2012. ⁴ Daniel C. Hallin & Paolo Mancini, Comparing Media Systems: Three Models of Media and Politics (Cambridge University Press, 2004). ⁵ Reported in the comparative literature on interpretive journalism; see Esser & Umbricht (2014) and related five-country comparisons. ⁶ Fabrizio Gilardi, Meysam Alizadeh & Maël Kubli, “ChatGPT Outperforms Crowd Workers for Text-Annotation Tasks,” PNAS, 120(30), 2023. ⁷ Petter Törnberg, “ChatGPT-4 Outperforms Experts and Crowd Workers in Annotating Political Twitter Messages with Zero-Shot Learning,” arXiv, 2023; see also his later comparison against supervised classifiers in Social Science Computer Review (2024).


