Glitchfield All articles
Culture

Dead Companies, Live Data: The Ghost Datasets Quietly Shaping Your Future

Glitchfield
Dead Companies, Live Data: The Ghost Datasets Quietly Shaping Your Future

Photo: smial (talk), FAL, via Wikimedia Commons

Picture a company that no longer exists making a prediction about your future. Not a metaphor — a literal event, happening right now, somewhere in the pipeline of a credit bureau or a hiring platform or an insurance algorithm. The company folded in 2019. The founders moved on. The app got a one-star farewell review from a user who just noticed it stopped working. But the data? The data is fine. The data has a whole new life.

This is what data persistence actually looks like in practice, and it's weirder and more consequential than most people realize.

The Startup Graveyard Has a Data Problem

The numbers on startup failure are grim and well-known. The majority of venture-backed companies don't make it past five years. What gets talked about less is what happens to the information those companies collected on their way down. User behavior logs. Purchase histories. Location data. Social graphs. Health app check-ins. Mood tracking entries that felt innocuous when you made them in 2017.

Under US bankruptcy law, customer data is frequently treated as a transferable asset — something that can be liquidated along with the office furniture and the domain name. The FTC has challenged some of these sales over the years, most famously blocking the transfer of RadioShack's customer data in 2015. But those interventions are the exception. Most of the time, when a startup dies, its data walks out the door quietly, absorbed into the inventories of data brokers, acquired by competitors, or folded into the training sets of whoever bought the IP.

Privacy policies almost always contain some version of the clause: In the event of a merger, acquisition, or sale of assets, user data may be transferred. Few people read it. Fewer still imagine it applies to a full collapse.

What "Outdated" Means When an Algorithm Is Using It

Here's where it gets philosophically strange. Data has a shelf life for humans — we intuitively understand that what someone did in 2015 might not tell you much about who they are in 2024. But machine learning models don't have that intuition built in unless someone explicitly engineers it.

When a model is trained on historical data, it learns the correlations that existed at the time of collection. If a defunct fitness app's dataset showed that users in certain zip codes had lower activity levels, and that correlated with some other outcome the model was trying to predict, the model will carry that signal forward indefinitely — even if those neighborhoods look completely different now, even if the app's user base was never representative to begin with, even if the whole dataset reflects a moment in time that has passed.

Researchers call this temporal drift or dataset shift, and it's one of the less glamorous problems in applied machine learning. It's also one of the most consequential. A model trained on pre-pandemic financial behavior is going to misread post-pandemic financial behavior. A hiring algorithm trained on data from a platform that attracted a specific demographic slice is going to replicate that slice's characteristics as a template for success.

The datasets don't know they're old. They just keep predicting.

The Brokers Who Keep the Lights On

The data broker industry in the US is enormous, largely unregulated at the federal level, and almost entirely invisible to the people whose information it trades. Companies like Acxiom, LexisNexis Risk Solutions, and hundreds of smaller players maintain profiles on most American adults — profiles assembled from sources that span decades, platforms, and companies that may have ceased to exist.

For these brokers, old data isn't necessarily bad data. It's inventory. The fact that it came from a company that no longer operates doesn't reduce its market value; in some cases it increases it, because the original collection constraints no longer apply. The app's terms of service said the data would only be used for personalized fitness recommendations. The app is dead. The terms died with it, at least in any practical enforcement sense.

Some of this data finds its way into AI training pipelines through channels that are difficult to trace. A model fine-tuned on a broker's dataset doesn't come with a bill of materials listing every source. The companies building those models often don't know the full provenance of what they're training on — or they know and don't ask too many questions.

When the Past Votes on Your Present

The real-world stakes here aren't abstract. Credit scoring, tenant screening, insurance pricing, and automated hiring tools all rely on data pipelines that frequently include third-party information of uncertain origin. When someone gets denied a loan or an apartment, the reason given is often generic enough that the specific data source is impossible to identify.

Consider what this means for people who have changed significantly since 2016 or 2018 — which is most people. Someone who went through financial hardship during COVID, rebuilt their credit, and is now applying for a mortgage might still be carrying the shadow of an old behavioral profile assembled from a now-defunct budgeting app that flagged their spending patterns during the worst months of their life. They can't dispute data they don't know exists. They can't request deletion from a company that no longer has a customer service department.

The Fair Credit Reporting Act gives consumers some rights around credit data specifically, but it doesn't cover most of the AI decision-making infrastructure that shapes hiring, housing, and insurance. And even where rights technically exist, exercising them requires knowing what data is being used — which the companies making decisions are rarely required to disclose.

Ghosts in the Training Set

There's a specific version of this problem that's becoming more visible as AI systems get more scrutinized: the presence of deprecated datasets in the lineage of current models. When researchers audit large language models or recommendation systems, they sometimes find that parts of the training data trace back to sources that no longer exist — forums that shut down, apps that got acquired, platforms that pivoted away from their original form.

Those datasets carry the cultural assumptions, demographic biases, and temporal quirks of their original context. A model trained partly on data from a social platform that was popular in 2014 will have internalized something about how people communicated and what they cared about in 2014. That's not inherently disqualifying, but it becomes a problem when nobody knows it's in there and nobody's accounting for it.

The push for AI transparency and model cards — documentation that explains what a model was trained on — is partly an attempt to surface this problem. But the documentation is voluntary, often incomplete, and rarely traces data provenance back far enough to catch the zombie datasets buried in the supply chain.

The Data Doesn't Care That You've Moved On

There's something almost poetic about the situation, in a grim sort of way. The internet has always had a memory problem — things that should fade don't, and things that should persist disappear. Dead company data is the same paradox running in reverse. The company is gone, but the record of who you were when you used it is still out there, still circulating, still occasionally surfacing to whisper something about you into a model that's making a decision you'll never fully understand.

The fix, if there is one, probably requires a combination of things: federal data broker regulation with actual teeth, stronger data deletion rights that survive corporate bankruptcy, mandatory provenance tracking in AI training pipelines, and a cultural shift in how we think about data as an asset versus data as a responsibility.

Until then, somewhere in a dataset with no living owner, a version of you from several years ago is still making predictions. Whether those predictions are accurate is almost beside the point. They're being trusted anyway.

All Articles

Related Articles

Undead Signal: The Underground Networks Keeping Cancelled Shows Alive Against All Odds

Undead Signal: The Underground Networks Keeping Cancelled Shows Alive Against All Odds

Ghost Platforms: The Dead Startups That Quietly Rewired the Internet

Still Online: The Faithful Keeping Abandoned Corners of the Web From Going Dark

Still Online: The Faithful Keeping Abandoned Corners of the Web From Going Dark