What problem does it solve?
Every earnings season, thousands of companies hold calls and file disclosures within a few weeks of each other. An analyst covering even a modest sector cannot read every transcript closely, let alone track how the same language choices played out across market cycles. Early text based signals in investing counted positive and negative words in a document to build a sentiment score, a method that misses tone, hedging, and how a sentence's meaning depends on the words around it.
The qualities that experienced fundamental analysts weigh, such as competitive position, pricing power, and whether an executive's tone on a call has shifted, have historically depended on individual judgment and were hard to measure systematically across a full coverage universe. A model that reads text the way a careful analyst does, at the scale of a whole index, gives portfolio managers a new, repeatable input, but only if its output is checked against experience rather than trusted on its own.
How does it work?
- Ingest text at scale. Earnings call transcripts, filings, and other public disclosures for the coverage universe are collected as they are published, alongside historical market reaction data for training and testing.
- Score with a fine tuned model or a structured prompt framework, not a general chatbot. A large language model is either fine tuned on a narrow, specific task, such as predicting the market's reaction to an earnings call, or run through a carefully engineered prompt and scoring framework, such as scoring a defined dimension of business quality, rather than used as an open ended assistant.
- Run it systematically. The model scores the full coverage universe on a set schedule (for example weekly), producing a consistent, comparable output per company rather than a one off read of a single transcript.
- Check it against what is already known. The new score is compared with existing quantitative measures and with analysts' own view, so the team can see where the two agree and, more importantly, where and why they diverge.
- Keep a person in charge of the decision. A portfolio manager or analyst defines the question, reviews the model's reasoning and the passages behind it, corrects or overrides results, and decides what, if anything, changes in the portfolio.
- Audience
- Employee facing
- Autonomy
- Copilot
- Adoption
- Emerging
- Channels
- Internal tools
What is it worth?
Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.
No public deployment has disclosed a measurable outcome yet.
Value drivers: Revenue growth, Speed and cycle time, Employee productivity.
Indicative value
An equity research team covering 500 companies
USD 50,000 to USD 400,000
Value of analyst time released from first pass transcript review per year
How this is calculated
Formula: analysts * transcriptsPerAnalystPerYear * hoursSavedPerTranscript * hourlyCost. The low scenario uses every low input, the high scenario every high input.
| Input | Low | High | Basis |
|---|---|---|---|
| Analysts and associates on the team analysts, people | 20 | 20 | The reference team. |
| Earnings call transcripts read per analyst per year transcriptsPerAnalystPerYear, transcripts per analyst per year | 100 | 200 | Editorial assumption for quarterly reporting across a broad coverage list, replace with your own coverage volume. |
| Hours of reading and note taking saved per transcript hoursSavedPerTranscript, hours per transcript | 0.25 | 0.5 | Editorial assumption, replace with your own time study; the model still needs a human read of the passages it flags. |
| Fully loaded analyst cost per hour hourlyCost, USD per hour | 100 | 200 | Editorial assumption, replace with your own fully loaded cost. |
What it leaves out: Values time saved on first pass reading only. It leaves out the value, positive or negative, of any investment decision the signal influences, the cost of building and running the models, and the analyst time spent reviewing the model's output instead.
Who already uses it?
2 public deployments, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.
T. Rowe Price
United States · Wealth and asset management · 2026
T. Rowe Price's Integrated Equity team built a large language model framework, engineered through an iterated prompt and scoring process rather than a fine tuned model, that assesses qualitative business quality by combining the model with the firm's own proprietary research and public information, then scores every company in the small and mid cap Russell 2500 Index on a weekly batch run. The team compared the new score with its existing quantitative quality score on the same universe: the two agreed on most companies but diverged on others, and T. Rowe Price used those disagreements to study a recent rally in lower quality stocks. A second case study on the same page applies a similarly built LLM prompt framework to score software companies on their resilience to AI disruption. T. Rowe Price calls the first analysis preliminary and flags look ahead bias and overfitting as open risks; for both analyses together it says they remain in the early stages of development and require additional validation.
No outcome disclosed.
BlackRock
United States · Wealth and asset management · 2024
BlackRock Systematic says it has used AI and machine learning in its investment process for nearly two decades and now uses large language models fine tuned on narrow, specific investment tasks rather than general purpose chatbots. One model is trained on more than 400,000 earnings call transcripts covering over 17,000 public firms, combined with two decades of historical market data, to learn an association between what is said on a call and the market's subsequent reaction. A separate tool, the Thematic Robot, blends the same kind of text analysis of transcripts and other sources with proprietary data so a portfolio manager can build an equity basket around an emerging market theme, with the manager defining the theme, reviewing the model's reasoning, and correcting or overriding its output.
No outcome disclosed.
How do you implement it?
A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.
Data you need
- A machine readable feed of earnings call transcripts and filings for the coverage universe
- Historical market reaction or outcome data to fine tune, or to build and test a prompt based scoring framework, against
- The team's existing quantitative factor or quality scores, to compare the new signal with
Systems to integrate
- Transcript and filings data provider
- Portfolio and research management system
- Quantitative factor and risk model platform
- Model validation and monitoring tooling
Complexity: High
Reading text is the easy part; making the score trustworthy is not. BlackRock fine tuned a proprietary model on more than 400,000 earnings call transcripts before relying on it in production. T. Rowe Price instead built a prompt engineered scoring framework on top of an LLM and calls its results preliminary, still pending additional out of sample testing. A lighter version, using a vendor's ready made transcript summaries or sentiment scores, is medium complexity but gives up some of the precision of either approach.
- 1
Pick one narrow, well defined task first
Choose a single, specific question, such as forecasting the market reaction to an earnings call or scoring one dimension of quality, rather than building a general purpose research chatbot. A narrow task is easier to fine tune, test, and trust.
- 2
Fine tune, or engineer a prompt framework, on your own outcome data
Either fine tune the model on historical transcripts paired with what actually happened afterwards, or build and iterate a structured prompt and scoring framework against the same data, so it learns the association for your universe rather than relying on a general purpose model's broad, unfocused training.
- 3
Compare against the existing quantitative view
Run the new score alongside current factor or quality scores on the same universe and look closely at the cases where they disagree; that is usually where the new signal earns, or loses, its keep.
- 4
Test out of sample before anyone relies on it
Hold back a period the model has not seen, walk the test forward in time, and check for look ahead bias and overfitting before the signal reaches a live portfolio process.
- 5
Build the review step in from day one
Give the portfolio manager or analyst the passages behind every score, not just the number, and make it easy to challenge, correct, or override the model's read.
Guardrails
- Every score is traceable to the transcript or filing passage that produced it
- No trade or position change executes from a raw model score without a portfolio manager's decision
- New signals are tested out of sample and checked for look ahead bias and overfitting before use
- The model is fine tuned or run through a structured scoring framework on a narrow, specific task rather than used as an open ended assistant for investment advice
KPIs to instrument
- Agreement and disagreement rate between the new score and the existing quantitative measure, reviewed by the team
- Out of sample performance of the signal, retested on a rolling basis as new data arrives
- Analyst hours spent on first pass transcript review, before and after
- Rate and reasons for analyst overrides of the model's score
Human in the loop
A portfolio manager or analyst defines the question the model answers, reviews the reasoning and source passages behind every score, and decides what, if anything, changes in the portfolio. The model flags evidence; it does not decide or execute a trade on its own.
Common failure modes
- Look ahead bias
- A model trained on data through the present can implicitly "know" what happened after the call it is scoring. Test strictly on a point in time history and walk forward only.
- Overfitting to the current market regime
- A prompt or scoring framework tuned on recent data may not hold up once conditions change. Retest out of sample as new data arrives and watch for a drop in agreement with the quantitative baseline.
- Treating a text score as settled fact
- A model's read of tone or quality can be wrong, or shaped by careful wording on the call. Keep the source passage next to every score and corroborate before sizing a position on it alone.
What are the risks and rules?
EU AI Act
Minimal risk
An internal research tool that scores companies for a firm's own portfolio managers is not listed in Annex III and is not a practice prohibited by Article 5. It carries no Article 50 transparency duty: those disclosure obligations, including the Article 50(2) marking duty on providers of systems that generate text, apply to content or interactions shown to a customer or the public, and this tool's output never leaves the firm's own research process; the portfolio manager, not the model, remains accountable for any resulting investment decision.
Rules that apply
Guidance
- SR 11-7, guidance on model risk management (Federal Reserve and OCC, North America). US supervisory expectations for validating, documenting, and monitoring quantitative models. Binding on Fed and OCC supervised banking organizations, including bank owned asset managers and broker dealers; other firms, such as the SEC registered advisers on this page, often use it as the reference standard for model validation but instead fall under Advisers Act compliance obligations.
Controls to put in place
- Model inventory entry with an accountable owner in the investment or quant research team
- Documented fine tuning data or prompt and scoring framework design, intended use, and known limitations for each model
- Scheduled out of sample retesting, with a defined threshold for retiring or retraining a signal
- Portfolio manager sign off recorded for any position materially influenced by a model score
Frequently asked questions
- Can an AI model really predict how a stock will react to an earnings call?
- BlackRock reports using a fine tuned model for exactly this. It says the large language models it uses for security analysis are trained and fine tuned on narrow, curated datasets for specific investment tasks, such as forecasting the market reaction following a corporate earnings call, rather than being general purpose chatbots. Treat any such signal as one input that is tested out of sample, not a forecast to size a position on by itself.
- How is this different from an AI research summarization assistant?
- A research summarization assistant condenses a firm's own published research and house view for advisors and analysts to use with clients. This use case instead extracts a new, systematic signal, such as a quality or resilience score, directly from company disclosures for a portfolio manager's own investment process, and it is judged by how well the signal performs out of sample, not by how faithfully it repeats a published view.
- What is the biggest risk in building this in house?
- Look ahead bias and overfitting. T. Rowe Price's own published case study on a large language model quality framework names both risks explicitly and describes its results as preliminary, pending further out of sample testing, which is the right level of caution for a signal like this.
How to cite this page
Blits.ai AI Use Case Library, "AI analysis of earnings calls and other text for investment signals", last verified 29 September 2026, https://www.blits.ai/ai-use-cases/earnings-call-and-text-analysis-for-investment-signals. Licensed under CC BY 4.0. Method: how we verify use cases.
Changelog
- 29 September 2026: First published