Google has acquired the data of failed US airline Spirit
Google’s purchase of Spirit Airlines’ corporate data out of bankruptcy for AI training is raising alarms over privacy, consent, and the treatment of personal information as a tradable asset. Commenters question how “deidentification” of millions of emails, call recordings, tickets, and internal documents can be meaningful when large tech firms can easily re-identify individuals by cross-referencing other datasets, and note that bankruptcy often overrides original usage consents. Others see the move as part of a broader trend: AI companies racing to acquire unique real‑world operational data to automate white‑collar work, while regulators struggle to update data protection rules fast enough.
Scope of the sale and legality
- Spirit’s internal digital exhaust (emails, Teams/SharePoint/OneDrive artifacts, ops data, support logs, tickets, etc.) was auctioned in bankruptcy and bought by Google, nominally to “improve AI services.”
- Commenters debate whether this should even be legal: many see it as a breach of customer and employee expectations, especially where consent was given for “training” humans, not LLMs.
- Some argue that in the US “who owns the computer owns the data” and bankruptcy courts prioritize creditor recovery, rewriting contracts unless specific statutes (e.g., GDPR-style) prevent it.
What data is included (disputed)
- Some quote coverage saying Google gets massive volumes of customer behavior data (calls, chats, email addresses, Wi‑Fi purchases, etc.).
- Others who read the court filing say “customer behavior” data is explicitly excluded from the Google purchase request and accuse the article of misreporting.
- There is mention of a “Deidentification Agent” (a third party picked by Google) and reference to California CCPA de-identification standards.
De-identification and re-identification
- Many are deeply skeptical that a trove of 100M+ emails, 30M calls, and rich operational logs can be meaningfully de-identified, especially when cross-linked with other datasets.
- Examples raised: prior AOL/Yahoo “anonymous” data sets, uniqueness from a few quasi-identifiers (gender, birthday, ZIP), and stylometry (writing-style fingerprinting).
- Some note that re-identifying de-identified data is nominally against big-tech policy and could be a fireable offense, but others point out incentives, weak enforcement, and “bounded distrust.”
Why Google would want this
- Seen as a gold mine of:
- Customer-service interactions for call/chatbot training.
- Rich, time-stamped corporate workflows across email, tickets, chats, and ops data to train “agent swarms” that simulate running a company.
- Unique internal style and patterns of real white‑collar work, beyond generic web text.
- A minority suggest it might inform blacklists, social-credit-like scoring, or fine-grained price discrimination, though this is framed as speculation and fear.
Privacy, regulation, and norms
- Numerous comparisons to GDPR: purpose-limited consent, non-transferable usage scopes, and meaningful fines vs. US’s business-friendly defaults.
- Some note GDPR is imperfect and evaded via “legitimate interests,” but still far better than nothing.
- Several argue that “de-identification” standards haven’t kept pace with ML capabilities and call for urgent new regulation.
Broader reactions
- Strong unease that “data as an asset” persists beyond a company’s life and is monetized in bankruptcy.
- Some resigned comments: most people already handed similar data to Gmail, Teams, etc.; data brokers and tracking are pervasive.
- Others focus on the dystopian direction: training models on human misery, loss of meaningful privacy for ordinary people, and AI as a tool for ever-tighter corporate control and extraction.