Tracer Research / Research paper 01

Pricing
Capability.

How do you price data when its real value is the capability it adds to an AI system? Spirit's selected $10 million bid raises the question: could the package have been underpriced? The paper shows what must be measured to answer it.

Adam Rida15 minute read

Read the article Read the paper
Abstract

Proprietary business records are becoming inputs for training, evaluating, and supervising AI systems. Their value depends on the task improvement they produce and the buyer's ability to turn that improvement into revenue or savings. This paper introduces Capability Alpha, a controlled estimate of a dataset's contribution to performance, and Price-Implied Uplift, the minimum commercial improvement required to recover an observed price. Spirit Aviation provides the worked case. Google's selected $10 million bid covers data and internally developed software. With 5% of annual Google Cloud revenue as an illustrative scope, it implies 0.236% persistent uplift over three years. Court approval remained pending, and Spirit's Capability Alpha remains unmeasured. Whether that amount understated the package's value to Google remains open.

Selected package bid
$10M
Data and internally developed software
Break-even uplift
0.236%
At 5% modeled Cloud scope over three years
Forward sensitivity
$99.6M
At 1% airline effect and 1% Google capture

The forward figure comes from chosen inputs. Evidence that value exceeded the bid would require a measured Capability Alpha and observed commercial capture.

The central idea, in three steps.

  1. 01
    Measure the technical gain.

    Capability Alpha isolates how much a proprietary dataset changes performance on a defined, economically relevant task.

  2. 02
    Translate it for a buyer.

    The same measured gain can be worth very different amounts depending on exposure, monetization, useful life, rights, and integration cost.

  3. 03
    Reverse an observed price.

    Price-Implied Uplift asks how much persistent commercial improvement would be needed to justify the price paid.

The Spirit illustration

A $10 million package price, applied to a scenario where 5% of annual Google Cloud revenue is the relevant base, implies a three-year break-even uplift of about 0.236%.

This threshold comes from scenario inputs. An experiment would still be required to measure Spirit's Capability Alpha, and the package includes data and software.

Data volume leaves the valuation question unanswered.

File, row, message, and token counts describe scale. Valuation begins with the task performance that changes because of the data: resolving a disruption, repairing software, forecasting demand, or completing another valuable task.

This distinction matters because data only creates economic value through a use. A billion disconnected records may be less useful than a smaller corpus that preserves what happened before a decision, which action was taken, which tools were used, and what happened next. For AI systems, that sequence can become a demonstration, a held-out evaluation case, a verifier, or post-training supervision.

What has changed?

Businesses have analyzed event logs and audit trails for decades. Today, workflow-linked records are being packaged for training, retrieval, evaluation, and verification across AI systems. One corpus may support several task families.

Earlier work already connects information and predictive improvement to economic value. This paper adapts that work to proprietary enterprise data transactions. It measures the corpus's incremental contribution, maps the result to a buyer's economics, and states the assumptions used to interpret an observed package price.

Which economically relevant task can this dataset improve, and by how much?

Measure task improvement first. Then connect it to the buyer's economics.

A dataset's price changes with the task, the experimental stack, and the buyer. Technical measurement belongs to the task and stack; economic translation belongs to the buyer.

01 / Input

Dataset

A legally usable corpus with defined structure, provenance, and rights.

02 / Measure

Capability Alpha

The task-level gain caused by access to the data under a controlled stack.

03 / Translate

Buyer economics

Exposure times monetization times persistence, adjusted for rights and cost.

04 / Output

Buyer value

Direct use value, a price-implied threshold, and supported strategic effects.

Technical question

Did the data cause a real improvement?

Capability Alpha answers this with a controlled evaluation. That result belongs to the experiment. Buyer size enters later, in the economic calculation.

Economic question

What is that improvement worth here?

Exposure, product fit, distribution, rights, useful life, and cost can make the same measured gain far more valuable to one buyer than another.

Capability Alpha is the gain from replacing reference data under an equal budget.

Fix the model-adaptation procedure and total resource budget. The treatment replaces a pre-specified amount of reference data with the same amount of proprietary data.

Equation 01Capability Alpha
CAD,t(k) = E[Yt(M1)] - E[Yt(M0)]

In words: expected held-out performance after replacing k units of reference data with proprietary data, minus the equal-budget reference-data control.

Dk

The pre-specified proprietary-data dose drawn from dataset D.

k

The amount of reference data replaced while the total adaptation budget stays fixed.

M0

The control model adapted on B units of reference data.

M1

The treatment model adapted on B-k reference units plus Dk.

Yt

The pre-specified metric on an independent held-out task evaluation.

๐”ผ

An average over matched training seeds and evaluation outcomes.

The reference design

Same base model, adaptation method, schedule, token or example budget, compute budget, and evaluation harness.

Control / M0 RB

B units of pre-specified reference data

Treatment / M1 RB-k + Dk

The same total budget, with provenance as the intervention

A small example

Suppose both arms use the same LoRA configuration and a 20-million-token adaptation budget. The control uses 20 million reference tokens. The treatment replaces 5 million of them with proprietary IT-support trajectories.

Across matched runs on 1,000 held-out incidents, control resolves 55.0% and treatment resolves 61.2%. Capability Alpha at a 5-million-token dose is +6.2 percentage points. That is a technical estimate under the declared stack, with no monetary assumption.

Data Shapley, Datamodels, TRAK, and LESS can help locate influential records or plan follow-up ablations. The paper treats them as diagnostics. The valuation input still comes from the direct matched experiment.

First check the capability gap

The paper first defines a diagnostic gap, Gt = Ht - Ft: the distance between a reference performance level and a practical baseline. A large gap shows room to improve. Evidence that dataset D contains the needed signal still requires a controlled experiment.

01

Pre-specify the task. Fix the target distribution, metric, and verifier before observing treatment results.

02

Hold out whole cases. Split by incident, customer, project, route, contract, or time period and screen for leakage.

03

Match the budget. Change the provenance of k units while model, optimizer, schedule, compute, and harness stay fixed.

04

Estimate uncertainty. Use paired seeds and intervals that reflect paired outcomes and run-to-run variation.

05

Keep contrary evidence. Retain negative effects and consider an equal-size irrelevant or shuffled placebo.

06

Test scale and transfer. Repeat across data doses, a second model or checkpoint, and adjacent tasks.

Why linked workflow data can matter

Context becomes useful when it remains connected to action and outcome.

ContextActionTool useOutcomeVerifier

Outcome linkage describes whether records preserve that sequence. Verifiability describes whether success can be judged by an objective, expert, or reproducible criterion. These corpus properties shape usefulness. They enter valuation through measured performance and usable rights.

Direct use value depends on the buyer's economic exposure.

A task-metric improvement acquires dollar value through its effect on revenue, cost, throughput, time, error, conversion, retention, or another economic outcome.

Equation 02Buyer-specific direct use value
Vbuse(D) =ฯD โˆ‘h=1T 1(1+r)h โˆ‘t (Eb,t,h ฮปb,t CAD,t dD,t,h) -Cb(D)

In words: add the buyer's expected task-level benefits over time, discount them to today, adjust for usable rights, then subtract the buyer's total costs.

CAD,t(k)

Measured gain

How much dose k from dataset D improves task t in the matched evaluation.

Eb,t,h x ฮปb,t

Economic exposure

The buyer's economic base and the mapping from one metric unit into commercial change.

dD,t,h / (1+r)h

Persistence and time

How much of the data advantage survives, then what a future benefit is worth today.

ฯD - Cb(D)

Rights and costs

Usable access, less acquisition, legal, deidentification, cleaning, training, and deployment cost.

What this explains

A dataset can have the same measured Capability Alpha for two firms and still have different values. A buyer with broader exposure, stronger distribution, better complements, or lower integration cost can turn the same technical gain into more money.

Price-Implied Uplift calculates the break-even improvement implied by a price.

PIU starts with the observed price and calculates the persistent revenue-equivalent uplift required to break even.

Equation 03Price-Implied Uplift
PIU=u*= P+Cb RscopemA(T,r)ฯ

In words: total package and integration cost, divided by the present-value profit base that the capability could affect.

The annuity inside the formula

A(T,r) = โˆ‘h=1T 1 / (1+r)h

For a three-year useful life at a 12% discount rate, A(T,r) is approximately 2.402.
P

Observed transaction price.

Cb

Buyer-specific integration and transaction costs.

Rscope

Annual revenue base plausibly exposed to the capability.

m

Operating margin applied to incremental revenue.

A(T,r)

Discounted annuity factor across useful life T at rate r.

ฯ

Rights-and-usability adjustment between zero and one.

Worked Spirit scenario
  1. $99.072B x 5% = $4.954B

    Annual Google Cloud revenue inside the illustrative scope.

  2. $4.954B x 35.6% = $1.763B

    Annual operating profit associated with that revenue base.

  3. $1.763B x 2.402 = $4.23B

    Three-year present-value profit base at a 12% discount rate.

  4. $10M / $4.23B = 0.236%

    Persistent revenue-equivalent uplift required to break even.

5.0% Illustrative economic base

The slider changes only the modeled business base. Aviation's actual share of Cloud revenue and Spirit's effect on that revenue remain unknown.

$99.072BAnnualized Cloud revenue
35.6%Operating margin proxy
3 yearsModeled useful life
12%Discount rate
Price-Implied Uplift 0.236%

$4.954B annual scoped revenue produces a $4.23B three-year present-value profit base.

The base calculation sets integration cost to zero and the rights adjustment to one. The agreement assigns third-party deidentification costs to Google, so the headline omits those costs.

Spirit provides a rare public transaction with the decisive experiment still outstanding.

The public record contains bid amounts and a detailed asset schedule. The corpus's contribution to any Google model or product remains unknown.

The transaction

Spirit Aviation Holdings designated Google as the successful bidder at $10 million, subject to court approval and required deidentification. Mercor was identified as the $7.5 million alternate bidder. A later $12.5 million micro1 offer arrived outside the initial auction process. It records willingness to pay and had not become a completed transaction in the paper's record.

The selected bid covers data and internally developed software, with no disclosed allocation between them. The PIU exercise assigns the full $10 million to the package. If the software has positive standalone value, that convention is an upper bound on the amount attributed to data. If the components are complementary, a clean allocation may not be economically meaningful.

Workplace100M

emails listed in the schedule

Collaboration500M

Microsoft Teams items

IT operations667K

information-technology tickets

Software516

repositories with roughly 30M lines of code

Commercial3.53B

booking-curve observations

Revenue7.51B

revenue transactions

Possible linkage 01

Pricing statedecisionbooking outcome

Possible linkage 02

IT ticketcode changedeployment result

The agreement requires deidentification while preserving referential integrity. Whether the real records are complete and usable enough to support these trajectories can only be established by inspection and controlled evaluation. Major consumer datasets, including customer profiles and Free Spirit member records, are excluded.

Could $10 million have been below the package's use value to Google?

The auction records bids from a specific process. The forward model calculates value under explicit improvement and capture assumptions. Its output is a sensitivity, and the Spirit evidence leaves the question unresolved.

Interactive sensitivity

Build a Spirit use-value scenario.

Choose an airline improvement and a Google capture rate. Airline revenue, Cloud margin, useful life, and discount rate remain fixed.

1.00% Applied to $1.165 trillion in 2026 global airline revenue as a scale assumption.
1.00% Assumed share of created airline value that becomes incremental Google revenue.
$1.165T2026 industry revenue
35.6%Cloud margin proxy
3 yearsAssumed life
12%Assumed discount rate
Illustrative present value of modeled operating profit $99.6M

$116.5M in annual incremental Google revenue produces this present value under the fixed margin, life, and discount assumptions.

Scenario PV รท selected $10M package bid10.0x
This tool shows how the stated assumptions change a present value. Spirit's Capability Alpha and commercial capture remain unmeasured. The forward present value excludes acquisition, deidentification, integration, and deployment costs, a rights adjustment, and strategic value.

How the range is constructed

The forward exercise begins with $1.165 trillion of global airline revenue. It then assumes an improvement in airline economics and a share of that created value captured by Google as incremental revenue. Applying the reported 35.6% Google Cloud operating margin, a three-year life, and a 12% discount rate produces the present values below.

The 1% airline effect and 0.5% to 2.0% capture rates are chosen scenario inputs. The table shows how present value changes across them; it carries no central forecast.

Present-value sensitivity to the two unobserved forward assumptions
Modeled airline improvement0.5% capture1.0% capture2.0% capture
0.25%$12.5M$24.9M$49.8M
0.50%$24.9M$49.8M$99.6M
1.00%$49.8M$99.6M$199.2M
2.00%$99.6M$199.2M$398.4M
The evidence supports

Observable bid signals, a transparent reverse-DCF threshold, and a conditional scale sensitivity.

Outside current evidence

Spirit's fair value, the share of the $10 million bid attributable to data, and any conclusion that the asset was underpriced.

Known

Publicly observed

The selected and alternate package bids, the disclosed asset schedule, and published Google Cloud segment inputs.

Assumed

Scenario choices

Revenue scope, margin as a proxy, useful life, discount rate, airline effect, and buyer capture rate.

Unknown

Evidence still needed

Actual Capability Alpha, post-deidentification usability, integration cost, data-software allocation, and realized commercial use.

Limited buyer participation can leave a gap between auction price and buyer-specific use value.

Legal access is one requirement. Monetization at scale also needs models, compute, post-training infrastructure, distribution, and commercial exposure.

01

Usable rights

Privacy rules, deidentification, license limits, and preserved linkage determine how much utility survives.

02

Capability gain

The corpus must improve performance on a task that still contains meaningful unsolved work.

03

Buyer exposure

The same gain is worth more when it can be applied across a larger economic base.

04

Complements

Compute, models, product surface, distribution, and integration capacity determine who can realize the gain.

Retain, license, or sell

The owner belongs in the same model.

Set the buyer index to the current owner and the framework produces internal retention value. A firm can rationally value its own data less than an outside buyer if it has less exposure, weaker commercialization channels, or higher integration cost.

Keep strategy separate

Preemption needs evidence of its own.

Restrictive rights may create value by denying a close rival access. A preemption estimate requires transaction rights and competitive evidence. Add strategic value only when that evidence is available.

Vbtotal(D) = Vbuse(D) + Sb(D)

What would weaken the thesis?

A useful valuation should expose its missing evidence. In the Spirit case, the decisive quantity still lacks a measurement.

Representative Spirit-derived tasks are already solved near the required quality, leaving little capability gap to close.

Controlled access or post-training produces no reliable held-out improvement after the rest of the stack is fixed.

Measured gains disappear across model families or fail to change real production outcomes.

Deidentification, weak linkage, restrictive rights, or high integration cost remove most practical utility.

Capability Alpha, after economic translation, remains below the PIU threshold under defensible scope assumptions.

Measurement

Capability Alpha is conditional on the evaluation stack. It requires controlled ablations, uncertainty intervals, contamination checks, and transfer tests.

Translation

The mapping from a task metric to revenue or cost may be nonlinear, and overlapping economic bases can create double counting.

Rights and cost

A package can lose value through privacy constraints, broken linkage, legal limits, deidentification burden, and costly integration.

Economic life

The advantage may decay as models improve, substitute data appears, or the underlying business process changes.

The complete research paper

Pricing Capability: A Framework for Valuing Proprietary Data Assets in the Age of AI

This article builds the intuition. The 19-page paper contains the literature review, formal model, matched-retraining protocol, sensitivity analysis, limitations, and bibliography.

Read the complete research paper.

Final PDF dated August 23, 2026. The source archive is being reconciled separately.

Open the paper

Sources and notes

Selected sources and method notes
  1. Pricing Capability: A Framework for Valuing Proprietary Data Assets in the Age of AI, Adam Rida, Tracer AI, Inc., August 23, 2026.
  2. Spirit Aviation Holdings auction results, asset purchase agreement, and data categories schedule, U.S. Bankruptcy Court filing, August 14, 2026.
  3. Alphabet Q2 2026 results, used for the annualized Google Cloud revenue and segment operating-margin inputs.
  4. IATA 2026 industry outlook, used for the airline-revenue scenario base.
  5. Google Cloud's SWISS case study and FLYR Labs case study, used only to establish that airline optimization can have measurable commercial consequences.
  6. These figures are conditional calculations from the August 23, 2026 paper. They carry no appraisal, investment recommendation, or claim about Google's intended use of the package. Spirit's Capability Alpha remains unmeasured.