mrkeyoor.com_
Tue 01 Sept 17:45 UTC
Tech6 min read

Google’s $10 Million Spirit Data Bid Includes Code and Emails

Google won a $10 million bid for Spirit Airlines’ deidentified corporate data, including code, email, finance, and flight records, pending court approval.

Google has won a $10 million auction for a large collection of Spirit Airlines’ corporate data, including 100 million emails, hundreds of millions of Microsoft Teams items, operational records, financial material, and the source history of 516 software repositories. The bid still requires approval from the bankruptcy court, but it shows how an airline’s internal digital record can become a saleable asset when the company itself shuts down.

The purchase matters beyond one failed carrier. AI companies are looking for material that is more structured and specialized than the public web, and a working company’s archives contain years of decisions, processes, conversations, and outcomes. Here, Google is buying not just documents but a detailed record of how a complex business operated.

Google’s bid was selected over a $7.5 million alternate offer from AI data company Mercor, according to an August 14 filing in the US Bankruptcy Court for the Southern District of New York. A hearing to approve the transaction is scheduled for August 19. Until the court signs off and the agreement’s conditions are met, Google is the successful bidder, not yet the final owner of the data.

What the $10 million bid covers

The asset schedule attached to the sale agreement is unusually specific. On the collaboration side, it lists 100 million corporate emails across 80,000 accounts, 17,082,644 OneDrive files, 20,577,677 SharePoint items, and 500 million Teams items. It also includes 667,563 ServiceNow IT tickets. The numbers appear to count records accumulated over time, not current employees or unique documents, so they should not be read as a workforce tally.

The code archive is substantial in its own right: 516 web, mobile, and back-end repositories totaling about 30 million lines, along with 43,170 pull requests and 372,585 commits. Issues, review discussions, CI/CD logs, test results, and code-coverage reports are also part of the package. Spirit is to provide the repositories as bare Git repositories or Git bundles and deliver metadata in JSON Lines format.

That combination may be more useful for software-focused AI work than source code alone. A repository shows the finished implementation; pull requests, bug reports, review threads, tests, and commit history show how people changed it, what failed, and which fixes were accepted. That is a richer record of engineering work, though the filing does not say Google will use it to train a coding model.

The included business data reaches much further. The schedule covers finance and accounting records, vendor and invoice material, forecasting files, board presentations, corporate-development documents, airline operations, crew pairing, maintenance, fuel, pricing, bookings, refunds, Wi-Fi sales, employee records, payroll documents, training records, tax material, and legal documents. The operational portion includes records for 763,391 Spirit flights, more than 1.2 million fuel slips, and 787,452 received parts. Legal forms, policy memos, litigation files, and contract revision histories are listed too, although legally privileged material is expressly excluded.

For Microsoft 365 material, the proposed transfer would retain the data in its native environment. Other systems are to be exported in machine-readable forms such as CSV, JSON, or SQL dumps. This is not a pile of scanned boxes. The agreement is designed to preserve links across datasets and make the archive usable as a connected technical corpus.

The customer call recordings are not in the deal

Early descriptions of the sale create a significant point of confusion. The Register reported that the purchase included more than 30 million customer-service call recordings, over 15 million chat sessions, and 13.7 million active marketing email addresses. Those figures do appear in the court filing’s inventory, but the adjacent purchase-request column marks those categories "Not Included."

The same exclusion applies to customer profiles, loyalty data, phone numbers, social-media material, surveys, website analytics, disability records, and customer complaints. The sale agreement also says the assets do not include Spirit’s customer list or information considered personal data under applicable privacy law.

The distinction is important. Google is bidding on a broad corporate and operational archive, not buying the customer database described in some coverage. Included systems may still contain personal information, especially employee, booking, payment, and legal records. That is why the agreement requires a separate deidentification process before transfer. But the schedule does not support the claim that the call recordings and chat transcripts themselves are being sold to Google.

Spirit retains the right to keep copies needed to wind down the business or comply with the law. It may also sell a separate customer list, including individual traveler spending aggregated by year, to companies in the hospitality or travel industries. That carve-out further shows that the customer list and Google’s data package are treated as different assets.

Google says the purpose is AI, but gives no product plan

Google said it bought the material to improve its AI services, according to The Register. The court documents do not identify a model, product, business unit, or training method, and Google has not laid out a public development plan in the material reviewed for this article. Claims that the archive is destined for a particular Gemini model, an aviation product, or an internal coding system would therefore be speculation.

Still, the contents explain the interest. Public text can teach a model language and general knowledge. A linked enterprise archive can expose the relationship between messages, tickets, code changes, operational events, financial decisions, and eventual outcomes. In principle, that could support research into agents that navigate business systems, models tuned to aviation operations, coding assistants trained on development history, or tools for analyzing corporate workflows. Those are possible uses suggested by the structure of the data, not disclosed Google plans.

Mercor’s alternate bid adds another signal. The company sells AI training and evaluation services, so two AI-oriented buyers independently assigned millions of dollars of value to the package. Yet the winning price alone does not establish how good the data is. The agreement sells the assets "as is" and disclaims warranties about accuracy, completeness, merchantability, and fitness for a particular purpose. Google also takes on the cost of preparing and deidentifying the collection, on top of the $10 million purchase price.

Deidentification is the central test

Before anything reaches Google, Spirit must send the selected data to one or more deidentification agents acceptable to the buyer. Those agents must remove or transform elements so the information cannot reasonably be linked to a particular consumer, while preserving referential integrity across the collection. The agreement calls for the California Consumer Privacy Act’s deidentification standard to be used for US consumer data even if that law would not otherwise apply. Health-related information must meet the relevant US health-privacy standard.

Google commits to keeping and using the material in deidentified form and not intentionally associating it with a person or household. It can transfer the data to third parties, but only if they are contractually bound to the same conditions. Privileged documents found after transfer must be returned or destroyed after Spirit gives notice.

Those are meaningful contractual controls, but they do not make the work trivial. Deidentifying a single table is different from cleaning an archive that connects email, employee systems, booking records, source control, legal files, and operational databases. Preserving the relationships that make such a corpus valuable can also preserve combinations of facts that point back to individuals. The filing sets a standard and requires certification; it does not demonstrate that the work has already been completed.

This transaction also puts a clear price on a question that usually stays hidden inside companies: what is the accumulated record of running a business worth once stripped from the business itself? Spirit’s aircraft, airport assets, and brand are easy to recognize as liquidation inventory. Its messages, code, process histories, and database exports are now inventory too. AI demand gives potential buyers a reason to pay for material that might once have been treated mainly as a storage and compliance burden.

The immediate things to watch are concrete: whether the bankruptcy judge approves the sale, what conditions appear in the final order, and how the deidentification certification is handled. After that, Google’s own disclosures will matter. Until it identifies a product or research result built from the archive, the strongest conclusion is narrower: a failed airline’s internal systems have become a $10 million AI-era data asset, and the legal and technical process for making that asset usable is only beginning.

We reviewed this

  1. flights — our honest review
  2. register — our honest review

Sources

  1. Notice of Auction Results and Scheduled Hearing for the Deidentified Data
  2. Google buys crashed airline Spirit’s data at auction, because AI