What a Driver’s License Breach Shows About Foreign Data-AI Threats

A major breach investigation is reportedly unfolding: An identity theft service started advertising digital scans of more than 153 million driver’s licenses from the United States and Canada on the dark web. The underlying service, as cybersecurity journalist Brian Krebs has reported, claims to have more than 10 million identification cards, 3 million travel documents or international IDs, and 579,000 medical cards. It vanished shortly after it was publicly discussed.

This breach is far more than another blow to data privacy and security. The incident is an excellent illustration of how the layering of breached data with AI analytic capabilities can let foreign adversaries build meticulous dossiers on virtually every American across sectors such as technology, defense, and retail. For American policymakers, counterintelligence professionals, and cyber defenders, considering how a foreign adversary such as the Chinese government could leverage breached datasets and AI tooling is essential to improving cybersecurity and data security measures in the near to medium term.

It can be tempting to dismiss the next hack of a database as “just another data breach.” Companies are constantly fending off attempted intrusions, and citizens have understandable alert fatigue from incident report after incident report. A common line of thinking is that because so much information is already out there, more will not make a difference.

But breaches do not exist in a vacuum. This breach reportedly exposed driver’s licenses, identification cards, travel documents, international IDs, medical cards, common access cards (although used as a general term, not necessarily in a U.S. defense context), residence cards, and employment authorizations. On its own, the information could already be valuable to a foreign adversary of the United States. However, the far greater danger—and the far bigger takeaway—lies in everything else a foreign adversary could combine it with, layering AI capabilities for data analysis on top.

Take China as an example: The Chinese government could theoretically purchase this data from the dark web and in doing so acquire the names, addresses, photos, travel identifiers, and other sensitive data of Americans. That scenario calls to mind building a searchable dataset of driver’s licenses or some other valuable, but fairly scoped, objective. And in an alternative world, that might be all the Chinese government would do with the data, if it were to acquire it.

Instead, consider how prolifically Chinese state actors penetrate U.S. databases. Over the last decade or so, the Chinese military hacked data broker Equifax and stole personal information concerning more than 145 million Americans. Sophisticated Chinese hackers stole over 78 million Americans’ health identification numbers, dates of birth, employment information, and other data from health insurer Anthem. Ministry of State Security personnel were reportedly behind the theft of more than 500 million Marriott guests’ data, including some people’s travel histories and passport numbers. Other Chinese hacks have spanned American medical records and airline data, while more recent operations have reportedly targeted telecommunications systems. And most notoriously, Chinese hackers broke into the Office of Personnel Management in 2015 and exfiltrated records on 22.1 million people, including security clearance holders. Collectively, that is an enormous volume of data, which, given the companies and government agency in question, likely came fairly well formatted and clearly labeled to boot.

Assembled together with the driver’s license data, the Chinese government’s picture from these acquisitions alone could look less like a searchable set of Americans’ driver’s license photos and more like a comprehensive set of dossiers about identities, finances, health conditions, travel, and communications. That could theoretically allow Chinese intelligence and security organs to identify Americans with embarrassing, stigmatized, or potentially vulnerability-revealing health conditions or treatments—such as a mental health condition or a sudden treatment for a sexually transmitted disease years into a marriage. It could theoretically allow Chinese operatives to not just look up people’s home addresses in driver’s license data (which they could also do through U.S. public records and other sources) but map faces against information on earnings, debts, and how often and where a person travels outside of the United States. Suddenly, the datasets look much more robust.

This national security problem set is not conceptually new. In fact, the “mosaic” concept in counterespionage refers to how “individual items of information can be combined with other data to provide a picture that discloses sensitive intelligence.” The notion that a nation-state could combine pieces of information it obtains from different sources is also part and parcel of espionage dating back centuries.

But the combination of data breaches, AI-assisted discovery of vulnerabilities to potentially enable more of those breaches, and sophisticated AI capabilities for data analysis brings this threat to a new level for the United States. Incessant data breaches and the subsequent sale of Americans’ data on the dark web provide the Chinese government with copious data sources in addition to what it steals itself. In addition to providing a fuller picture of a person’s or organization’s characteristics and activities, combining datasets has long made it generally easier to strip away any pseudonymization or “anonymization” protections, reidentify specific people contained within the data, and trace their activities across previously obscured datasets. Modern AI capabilities make this even easier: The Chinese government could certainly run models and algorithms over datasets to identify hidden patterns, uncover novel insights, and further efforts at data reidentification.

With these data and technology trends in hand, a breach of a large driver’s license database is no longer just one breach. For any nation-state that perpetrates the intrusion (although there is no evidence of that here) or gets access to the data after the fact, the right dataset can effectively multiply in value to the adversary—and compound damage imparted on the United States. The insights it could provide an adversary such as Beijing for intelligence operations, military planning, counterespionage, and disruptive cyber activity are augmented when datasets can be combined, reidentified, and subject to pattern and anomaly analysis with advanced AI capabilities. Whether used for facial recognition on license photos or correlation of names with records from hotels and airlines, the intelligence opportunities for an adversary—and harms for the United States—could also be retrospective, real-time, or forward-looking.

Although the Chinese government was used here as an example, similar threats come from other foreign adversaries, such as the Russian government. Its expansive web of cyber actors, sophisticated espionage capabilities, and potential access to breached and leaked data, among other factors, likewise affords it the opportunity to build extensive profiles of Americans.

Therein lies the real takeaway. U.S. policymakers, planners, and cyber defenders must ensure going forward that they do not try to prevent, mitigate, and investigate data breaches from nation-states (or breaches that could be exploited by them) in isolation. To prioritize defenses of various datasets, policymakers, planners, and cyber defenders must consider what gaps the adversaries most interested in stealing them might be seeking to fill. When investigating data breaches after the fact, post-incident autopsies from a counterintelligence perspective should always account for how dataset combination, reidentification, and AI analysis can make unsuspecting datasets particularly sensitive. And at a strategic level, the U.S. government should build on discussions in the past several administrations about foreign adversaries and big data to increasingly treat corporate-held and personal data as strategic assets that require serious protections.

As the FBI reportedly investigates the sale of breached license data, it may very well release information about the perpetrators. But the real risks have less to do with smiling driver’s license photos and more with how U.S. adversaries may use them to make enormous dossiers even thicker.

Justin Sherman is a senior associate (non-resident) with the Intelligence, National Security, and Technology Program at the Center for Strategic & International Studies in Washington, D.C.

Image
Justin Sherman
Senior Associate (Non-resident), Intelligence, National Security, and Technology Program