The package
The records, software, and legal rights actually included in the deal.
Tracer Research / Research paper 01
In Spirit Airlines' bankruptcy proceedings, a package of deidentified operational records and internally developed software was proposed for sale. Google is the selected bidder at $10 million, subject to court approval. That gives the public a price. What remains unknown is whether the data improves an AI system on a useful task and what any measured gain would be worth to Google. Could the package have been worth more?
Companies accumulate records of real decisions: customer demand, pricing changes, operational disruptions, IT incidents, and what happened next. Those records can train or test AI systems, yet their value is hard to observe because most transactions and controlled tests are private.
Spirit creates a public case because the public record describes the package and provides price signals from two auction bids and a later reported offer. The paper asks two questions. Capability Alpha estimates the data's effect on a valuable task through controlled tests. Price-Implied Uplift finds the smallest persistent commercial improvement needed to recover an observed price.
With an illustrative scope equal to 5% of annual Google Cloud revenue, a 0.236% annual revenue-equivalent uplift sustained for three years would recover the selected bid. The data's actual task effect remains unmeasured, so underpricing remains an open question.
The $10 million bid is observed. The percentages and value range come from stated assumptions. The compared contracts differ, and the data's performance contribution is unmeasured. Underpricing remains an open question.
In brief
AI companies are beginning to buy records of real work. Sellers have little public evidence for what those records are worth. A price shows what one buyer offered. A controlled test is still needed to show what the data changes.
Hold the system, task, and test conditions fixed. Compare repeated evaluations with and without the dataset. The estimated attributable difference is Capability Alpha.
Estimate how many decisions or workflows the improvement can affect, what each improvement earns or saves, how long it lasts, what uses the contract permits, and what integration will cost.
Start from the known price and calculate the smallest persistent commercial improvement needed to recover it. The paper calls this Price-Implied Uplift.
Suppose the package affects a business area equal to 5% of annual Google Cloud revenue. With the paper's margin, three-year life, and discount assumptions, a 0.236% annual revenue-equivalent uplift in that modeled scope, sustained for three years, would recover the $10 million bid.
The 5% scope is a modeling choice. Google's specific deployment surface and use mix, the data's actual effect on task performance, and the package's data-software allocation remain unknown.01 / The question
Spirit's schedule lists 100 million emails, 500 million Microsoft Teams items, 667,000 IT tickets, 516 code repositories, and billions of booking and revenue records. Those counts establish scale. They cannot tell a buyer whether the records make an AI system more accurate, faster, safer, or cheaper on a useful task.
Linked records preserve the context around a decision. An IT ticket connected to a code change and deployment result captures a problem, an action, and an outcome. A pricing decision linked to later bookings does the same. Those sequences can become training examples, test cases, or signals that an AI answer or action worked.
Connections between records matter. A smaller set of coherent histories can be more useful than a much larger archive of fragments. Valuation therefore starts with a defined task and a controlled measurement of what changes when the AI system gains access to the data.
Businesses have analyzed event logs and audit trails for decades. AI systems create more possible uses for the same records. They can learn from linked cases, retrieve relevant examples while working, and be tested against outcomes captured in operations. Each proposed use still needs a defined task and a measurable result.
Which economically relevant task can this dataset improve, and by how much?
02 / The method
Run matched tests on a stated task, changing access to the dataset while keeping the rest of the setup fixed. Then apply the measured result to the buyer's scale, products, permitted uses, useful life, and costs. This keeps experimental evidence separate from commercial assumptions.
The records, software, and legal rights actually included in the deal.
The performance difference between matched runs with and without the data.
The activity the improvement can affect, the revenue or savings per unit, how long it lasts, and the costs of using it.
Estimate what the use could be worth to that buyer, or calculate the minimum gain needed to recover a known price.
The answer applies to the tested system, task, and metric. Buyer size enters later, in the economic calculation.
The answer depends on operating scale, how each unit of improvement changes revenue or cost, time, contract rights, and implementation costs.
03 / The technical measure
Compare the same AI system on the same test twice, once with the data and once without it. Keep the model, compute, prompts, tools, and scoring procedure fixed. If data and software change together, test the combinations separately.
In words: expected held-out performance with access to dataset D, minus expected performance without it, while the remaining experimental stack stays fixed.
DThe proprietary dataset being evaluated.
tThe pre-specified, economically relevant task family.
M0The controlled model or system without access to dataset D.
MDThe controlled model or system with access to dataset D.
YtThe pre-specified metric on an independent held-out task evaluation.
𝔼An average over evaluation outcomes and experimental variation.
Same model family, compute budget, tools, harness, target distribution, metric, and evaluation setup.
The declared baseline stack without access to the proprietary dataset
Access to the proprietary dataset with the remaining stack held fixed
To make the definition concrete, imagine two model or system configurations using the same model family, tools, compute budget, and evaluation harness. The control has no access to proprietary IT-support trajectories. The treatment receives them.
Across controlled runs on 1,000 held-out incidents, control resolves 55.0% and treatment resolves 61.2%. Capability Alpha is +6.2 percentage points. That is a technical estimate under the declared stack, with no monetary assumption.
Data Shapley, Datamodels, TRAK, and LESS can help locate influential records or plan follow-up ablations. The economic valuation still requires a controlled estimate of the corpus's causal contribution.
Frontier Deficit is the distance between a reference performance level and current frontier performance. It shows how much valuable capability remains missing. Whether dataset D can close that gap still requires controlled evaluation.
Pre-specify the task. Fix the target distribution, metric, and verifier before observing treatment results.
Hold out whole cases. Split by incident, customer, project, route, contract, or time period and screen for leakage.
Hold the stack fixed. Keep the model family, compute budget, tools, harness, and evaluation setup constant while varying access to the dataset.
Estimate uncertainty. Use paired seeds and intervals that reflect paired outcomes and run-to-run variation.
Separate moving parts. Use a factorial design when data, tools, rewards, or harness components change together.
Test scale and transfer. Repeat across data fractions, model families or checkpoints, and adjacent tasks.
Verifier Density describes how much of the corpus supports objective or structured judgment of success. Outcome Linkage describes how reliably records connect context, action, intermediate state, and outcome. These properties shape usefulness through measured performance and usable rights.
04 / From performance to money
A task-metric improvement acquires dollar value through its effect on revenue, cost, throughput, time, error, conversion, retention, or another economic outcome.
In words: add the buyer's expected task-level benefits over time, discount them to today, adjust for usable rights, then subtract acquisition-adjacent costs.
CAD,tThe causal improvement in economically relevant model or system performance attributable to dataset D on task t.
Eb,t,h x λb,tThe buyer's economic base and the mapping from one metric unit into commercial change.
dD,t,h / (1+r)hHow much of the data advantage survives, then what a future benefit is worth today.
ρD - CDUsable access, less acquisition-adjacent legal, deidentification, engineering, cleaning, training, and integration costs.
A dataset can have the same measured Capability Alpha for two firms and still have different values. A buyer with broader exposure, stronger distribution, better complements, or lower integration cost can turn the same technical gain into more money.
05 / Reverse the price
PIU starts with the observed transaction price and recovers the persistent attributable commercial improvement required for the transaction to break even.
In words: observed transaction price plus acquisition-adjacent costs, divided by the present-value profit base that the capability could affect.
A(T,r) = ∑h=1T 1 / (1+r)h
For a three-year useful life at a 12% discount rate, A(T,r) is approximately 2.402.PObserved transaction price.
CAcquisition-adjacent legal, deidentification, engineering, cleaning, training, and integration costs.
RscopeAnnual revenue base plausibly exposed to the capability.
mOperating margin applied to incremental revenue.
A(T,r)Discounted annuity factor across useful life T at rate r.
ρRights-and-usability adjustment between zero and one.
Annual Google Cloud revenue inside the illustrative scope.
Annual operating profit associated with that revenue base.
Three-year present-value profit base at a 12% discount rate.
Persistent revenue-equivalent uplift required to break even.
The slider changes only the modeled business base. Aviation's actual share of Cloud revenue and Spirit's effect on that revenue remain unknown.
$4.954B annual scoped revenue produces a $4.23B three-year present-value profit base.
The base calculation sets integration cost to zero and the rights adjustment to one. The agreement assigns third-party deidentification costs to Google, so the headline omits those costs.06 / The case
The public record contains bid amounts and a detailed asset schedule. The corpus's contribution to any Google model or product remains unknown.
Spirit Aviation Holdings designated Google as the successful bidder at $10 million, subject to court approval and required deidentification. Mercor was identified as the $7.5 million alternate bidder. A later $12.5 million micro1 offer arrived outside the initial auction process. It records willingness to pay and had not become a completed transaction in the paper's record.
The selected bid covers data and internally developed software, with no disclosed allocation between them. The PIU exercise assigns the full $10 million to the package. If the software has positive standalone value, that convention is an upper bound on the amount attributed to data. If the components are complementary, a clean allocation may not be economically meaningful.
emails listed in the schedule
Microsoft Teams items
information-technology tickets
repositories with roughly 30M lines of code
booking-curve observations
revenue transactions
Pricing state→decision→booking outcome
IT ticket→code change→deployment result
07 / Price and value
The bids are market evidence. The forward range is a conditional buyer-specific use-value scenario. It shows which assumptions produce values of approximately $50 million to $200 million and makes no fair-value claim.
Choose an airline improvement and a Google capture rate. Airline revenue, Cloud margin, useful life, and discount rate remain fixed.
$116.5M in annual incremental Google revenue produces this present value under the fixed margin, life, and discount assumptions.
The forward exercise begins with $1.165 trillion of global airline revenue. It then assumes an improvement in airline economics and a share of that created value captured by Google as incremental revenue. Applying the reported 35.6% Google Cloud operating margin, a three-year life, and a 12% discount rate produces the present values below.
The 1% airline effect and 0.5% to 2.0% capture rates are chosen scenario inputs. The table shows how present value changes across them; it carries no central forecast.
| Modeled airline improvement | 0.5% capture | 1.0% capture | 2.0% capture |
|---|---|---|---|
| 0.25% | $12.5M | $24.9M | $49.8M |
| 0.50% | $24.9M | $49.8M | $99.6M |
| 1.00% | $49.8M | $99.6M | $199.2M |
| 2.00% | $99.6M | $199.2M | $398.4M |
Observable bid signals, a transparent reverse-DCF threshold, and a conditional scale sensitivity.
Spirit's fair value, the share of the $10 million bid attributable to data, and any conclusion that the asset was underpriced.
The selected and alternate package bids, the disclosed asset schedule, and published Google Cloud segment inputs.
Revenue scope, margin as a proxy, useful life, discount rate, airline effect, and buyer capture rate.
Actual Capability Alpha, post-deidentification usability, integration cost, data-software allocation, and realized commercial use.
08 / Cross-transaction reference
The paper applies Price-Implied Uplift to four additional AI data and content agreements. Three involve mature public-company buyers and support a standardized 5% scope comparison.
F is the annual fee, s is the modeled economic scope, and OI is buyer operating income used as a standardized proxy.
Standardized 5% break-even thresholds across three reported recurring agreements.
| Transaction | Reported consideration | PIU, stated 5% scope |
|---|---|---|
| Spirit / Google | $10M one-time package | 0.236%* |
| New York Times / Amazon | Reported $20M-$25M per year | 0.583%-0.729% |
| News Corp / Meta | Reported up to $50M per year | ≤1.201% |
| Reddit / Google | Reported about $60M per year | 1.424% |
* Spirit uses the paper's case-specific Google Cloud denominator and three-year discounted cash flow. The recurring-license rows use 5% of each buyer's pre-deal consolidated operating-income base as a standardized proxy.
Under the stated assumptions, Spirit has the lowest break-even hurdle in this small panel. A relatively small persistent contribution could justify the selected price. Whether the package was underpriced still depends on its unmeasured Capability Alpha.
Rights, duration, refresh, intended use, asset composition, and economic bases vary across the agreements. Spirit's Capability Alpha has never been measured.
The News Corp / OpenAI agreement is the fourth case. Reported consideration includes cash and OpenAI credits, and the public record provides no suitable positive operating-margin base. The paper reports a sensitivity because a point PIU would require assumptions that public evidence cannot support.
09 / Who can use it
Legal access is one requirement. Monetization at scale also needs models, compute, post-training infrastructure, distribution, and commercial exposure.
Privacy rules, deidentification, license limits, and preserved linkage determine how much utility survives.
The corpus must improve performance on a task that still contains meaningful unsolved work.
The same gain is worth more when it can be applied across a larger economic base.
Compute, models, product surface, distribution, and integration capacity determine who can realize the gain.
Set the buyer index to the current owner and the framework produces internal retention value. A firm can rationally value its own data less than an outside buyer if it has less exposure, weaker commercialization channels, or higher integration cost.
Restrictive rights may create value by denying a close rival access. A preemption estimate requires transaction rights and competitive evidence. Add strategic value only when that evidence is available.
Vbtotal(D) = Vbuse(D) + Sb(D)
10 / Limits and falsification
The decisive measurement is still missing: Spirit's actual Capability Alpha.
Representative Spirit-derived tasks are already solved near the required quality, leaving little Frontier Deficit to close.
Controlled access or post-training produces no reliable held-out improvement after the rest of the stack is fixed.
Measured gains disappear across model families or fail to change real production outcomes.
Deidentification, weak linkage, restrictive rights, or high integration cost remove most practical utility.
Capability Alpha, after economic translation, remains below the PIU threshold under defensible scope assumptions.
Capability Alpha is conditional on the evaluation stack. It requires controlled ablations, uncertainty intervals, contamination checks, and transfer tests.
The mapping from a task metric to revenue or cost may be nonlinear, and overlapping economic bases can create double counting.
A package can lose value through privacy constraints, broken linkage, legal limits, deidentification burden, and costly integration.
The advantage may decay as models improve, substitute data appears, or the underlying business process changes.
This article builds the intuition. The 21-page paper contains the literature review, formal model, controlled-evaluation design, Spirit sensitivity analysis, public-transaction calculations, limitations, and bibliography.
Current 21-page final submission, dated August 25, 2026.
Open the paper