Separating real AI efficiency from AI hype in earnings calls
- AI is mentioned on nearly every call now, so the mention itself carries no information. The signal is whether the claim is specific and quantified, and whether it shows up in the financials.
- Real efficiency claims name a function, a number, and a timeline; hype uses 'AI-powered' and 'leveraging AI' with nothing measurable attached. That distinction is scoreable across a whole universe.
- The check that separates the two is the financials: a genuine efficiency gain should be visible in margins, operating expense, or headcount growing slower than revenue. A claim with no trace in the numbers is narrative.
- One quarter is not proof. A credible AI-efficiency story is repeated across quarters with follow-through; a buzzword appears once when it is fashionable and disappears.
There is a stretch of every earnings season now where it feels like the same sentence is being read by three hundred different CEOs: some version of "we are leveraging AI to drive efficiency across the organization." The mention has become free. Because everyone says it, the fact that a company mentioned AI tells you nothing - and a keyword screen that flags "AI" returns the entire market. The screen worth building does something harder: it separates the companies making a specific, quantified, checkable efficiency claim from the ones dropping a buzzword to protect the multiple. Here is how.
The mention is not the signal
Start by throwing out mention count. When 90 percent of a universe references AI on the call, ranking by how often a company said "AI" sorts for talkativeness, not substance. The information is entirely in the character of the claim, which falls cleanly into two buckets: substantive and promotional. Substantive claims name a function, attach a number, and give a timeframe - "we reduced support headcount needs by using an AI agent that now handles 40 percent of tickets." Promotional claims are grammatically confident and empirically empty - "our AI-powered platform is transforming how we operate." That distinction is scoreable, and it is the whole game.
What a real efficiency claim looks like
Three features separate a claim you can weigh from one you cannot:
| Feature | Substantive | Hype |
|---|---|---|
| Specificity | Names the function or process affected | "Across the organization" |
| Quantification | A number: cost, headcount, cycle time | No number attached |
| Follow-through | Updated and built on across quarters | Appears once, when fashionable |
Scoring each AI passage against these three, the same way on every call, turns a universe of identical-sounding mentions into a ranking of credibility. The mechanics of pulling and scoring those passages consistently are the same as any transcript workflow, covered in analyzing earnings call transcripts at scale.
The financials are the lie detector
A claim about efficiency is a claim about the cost structure, and the cost structure is measured. So the decisive check is not in the transcript at all - it is whether the claim shows up in the numbers. Genuine AI-driven efficiency should be visible as expanding margins, operating expense falling as a share of revenue, or headcount growing slower than the business. A company narrating heavily about AI efficiency while margins are flat and the headcount line keeps climbing is describing a future, not a result. Pairing the verbal claim with the reported results is the same discipline as cross-checking earnings-call claims against filings - the gap between what was said and what the numbers show is the signal.
Why this is a management-quality read too
How a management team talks about AI is a tell about how they talk about everything. A team that makes specific, quantified, falsifiable claims and comes back next quarter with an update is behaving one way; a team that reaches for the fashionable buzzword and never quantifies it is behaving another. That places this squarely alongside the broader management-quality read, where specificity and follow-through are the traits worth scoring across the board.
What breaks
- Keyword-only screening.Counting "AI" mentions returns the whole market. The classification of the claim, not its presence, is what matters.
- Skipping the financial check. Scoring the language alone rewards good storytelling. The claim has to be checked against margins and headcount, or you are grading rhetoric.
- One-quarter snapshots. A single fashionable mention is not a strategy. Credibility comes from repetition with follow-through, which needs a multi-quarter view.
Running it as a pipeline
In Cutonce this is a saved pipeline: a transcript and filings node pulls the AI-related passages for every name in the universe, an AI node classifies each as substantive or promotional and extracts any quantified claim, and a step cross-references those claims against reported margin and headcount trends. The output is a ranked table separating the companies whose AI-efficiency story is backed by the numbers from the ones narrating, delivered to a sheet or Slack and rerun each quarter so the follow-through is visible.
The point is not to be cynical about AI - some of these claims are real and matter a great deal. The point is that the word has stopped carrying information, so you need a screen that reads the claim, not the keyword, and checks it against the results.
Note: this is not investment advice. Scoring AI claims produces a credibility ranking, not a valuation, and models can misread both the language and the financial context. Verify any name against its actual transcript and results before acting on it.
Frequently asked
How can you tell a real AI efficiency claim from hype? Specificity and verification. A real claim names what the AI is doing (which function or process), attaches a number (cost saved, headcount avoided, cycle time cut) and a timeframe, and is consistent with the financials - margins expanding or headcount growing slower than revenue. Hype uses vague phrasing ('AI-powered', 'leveraging AI') with nothing measurable, and leaves no trace in the numbers. Screening for the difference means scoring each AI mention on specificity and then cross-checking it against the actual results.
Can you screen earnings calls for AI mentions at scale? Yes, but a keyword screen alone is nearly useless now because almost every company mentions AI. The useful version pulls the AI-related passages from every transcript and filing, classifies each as substantive (specific, quantified) or promotional (vague, unquantified), and ranks companies by how credible and financially-backed their AI claims are. The classification, not the keyword count, is what carries the signal.
Why check AI claims against the financials? Because a genuine efficiency gain has to show up somewhere - in gross or operating margin, in operating expense as a share of revenue, or in headcount that grows slower than the business. A company that talks extensively about AI-driven efficiency while margins are flat and headcount is climbing is telling a story the numbers do not support. Cross-checking the claim against the results is what turns a management assertion into something you can weigh.
Is a strong AI efficiency claim a buy signal? No. It is a data point about management credibility and operating leverage, not a valuation. A company can be executing genuinely on AI efficiency and still be expensive, and a credible claim is only one input among many. Treat the screen as a way to separate the companies actually changing their cost structure from the ones narrating, then do the real work on the ones that pass.