TL;DR: Synthetic identity fraud detection is hard because the customer passes every identity check by design — the Social Security number is real, the name and date of birth cohere, the address is deliverable. The Federal Reserve found that fraud models built for traditional identity fraud fail to flag 85 to 95 percent of synthetic identity applicants. The losses that follow are usually written off as credit losses, which means the model never learns the pattern and the same identity can be rebuilt and used again.
What a Synthetic Identity Is
A synthetic identity is a fabricated person assembled from a combination of real and fake personally identifiable information. The Federal Reserve's industry definition, produced by a convened focus group of fraud experts and published in 2021, describes it as "the use of a combination of personally identifiable information (PII) to fabricate a person or entity in order to commit a dishonest act for personal or financial gain."
The distinction that matters operationally is between synthetic identity fraud and traditional identity theft. In identity theft, a real person eventually notices and calls the bank. In synthetic identity fraud there is no victim to complain, because there is no person. Detection has to be proactive or it does not happen at all.
The Fed's toolkit describes three construction methods: identity compilation, which combines real and fake PII and is the dominant form; identity manipulation, which slightly modifies real PII; and identity fabrication, which invents everything. Fraudsters preferentially use the Social Security numbers of people who have no active credit profile — children, the elderly, and the incarcerated — because nobody will check for years.
How the Identity Gets Built and Aged

The lifecycle runs through five stages, and the third one is where the financial system does the fraudster's work for them.
Step three is the structural flaw. As the Federal Reserve's October 2019 white paper puts it, when a fraudster first uses a synthetic identity to apply for credit, "the bureau creates a credit profile for the synthetic identity, which helps legitimize its identity even when credit is denied." A declined application still manufactures a bureau record. The system lenders consult to verify legitimacy is the same system fraudsters use to manufacture it.
Piggybacking accelerates the process. Adding the synthetic as an authorized user on a real, well-aged account transfers the primary user's established history, which builds a credible score in weeks rather than years.
The Fed's own worked example shows what the endgame looks like from the lender's side. An application presents a 750 FICO score, an oldest tradeline of 20 years — though it is an authorized user tradeline — only unsecured lines reported, eight recent inquiries, and a stated income of $125,000. The account is approved at $20,000. Three years of perfect payments doubles the line to $40,000. The $40,000 is maxed and never repaid. Because there were no indicators of fraud, the loss is charged off as a credit loss and never reported as fraud.
Why the Losses Are Invisible
That misclassification is the deepest problem in this typology, and it compounds in three directions.
Fraud models trained on labeled fraud never see these cases, so they never learn the pattern. That is a large part of why the Federal Reserve's white paper reported that models built to predict traditional identity fraud "did not flag 85% to 95% of potential synthetic identity fraud applicants." The training data is systematically missing the thing the model is meant to catch.
The delinquency is also purged by the bureaus after seven years, at which point the same synthetic can be rebuilt and busted out again.
And the aggregate loss number becomes unknowable. FinCEN's 2021 identity-related trend analysis recorded roughly 3,000 BSA reports totaling $182 million for the synthetic identity typology. Industry estimates for the same period run into the tens of billions. The Federal Reserve Bank of Boston, citing FiVerity, put losses at more than $35 billion in 2023. TransUnion measured $3.3 billion in U.S. lender exposure for the year ending 2024. The gap between the measured federal figure and the industry estimates is more than two orders of magnitude, and that gap is itself the finding.
The Risk Curve Runs Backwards

The most counterintuitive fact in synthetic identity detection is where the risk concentrates. TransUnion's analysis of 90-day delinquency by credit tier found synthetics in subprime performing almost identically to real subprime borrowers, a lift of only 1.1x. In prime plus, the lift was 3.6x. In super prime, synthetics were 12.5 times more likely to go 90 days delinquent than real customers — 7.5 percent against 0.6 percent.
The reason is that fraudsters invest heavily in building pristine profiles precisely to access higher limits. Your riskiest synthetics are hiding in your best-performing segment. Any program that concentrates fraud scrutiny on low credit tiers is looking exactly where synthetics are least distinguishable.
Payment behavior is not a signal either. Cultivated synthetics are model customers by design. Data cited in the Fed's white paper puts the average charge-off rate for likely synthetic identities in a lending portfolio below 30 percent, which implies that around 70 percent are sitting quietly in the portfolio exhibiting typical consumer payment patterns.
What Actually Detects Them
The Fed's toolkit is direct that no single indicator signifies a synthetic identity. The signals that work are relational and cross-account.
At onboarding, the indicators worth building on are a mismatch between the applicant's age and the duration of their credit history, a credit profile consisting primarily of authorized user tradelines, a Social Security number issued after 2011 paired with a date of birth well before it, high unsecured debt relative to file depth, and a high number of recent inquiries. TransUnion's 2025 research found that the absence of public records is unusually predictive — no known relatives and no motor vehicle registration appear in 30 to 50 percent of synthetic identities and increase the likelihood of a profile being synthetic by up to seven times.
The Fed calls cross-account matching the biggest indicator: matching contact information across identities, the same Social Security number appearing under multiple names, shared device fingerprints, and identical digital footprints. Applications from data center IP addresses, VPNs, or proxies rather than residential connections, and identity elements that show no natural time-built linkage to one another, both belong in the same category.
In-life, the sharpest single signal the Fed names is payment source analysis: whether payments to a credit card came from the same external account remitting payments to other cards under different names. Cross-product sweeps matter here too. Synthetics distribute across checking, savings, cards, loans, and merchant services, and a fraud review scoped to one product line will not see the cluster. The same relationship-level view underpins perpetual KYC, which is why institutions that have built one usually find the other cheaper.
Verification tooling has a floor. The Social Security Administration's eCBSV service confirms that a name, date of birth, and Social Security number match SSA records and returns a death indicator. That tells you the tuple is real. It does not tell you a person is attached to it.
Generative AI Changed the Economics
The Federal Reserve Bank of Boston's April 2025 assessment is that generative AI has made this materially worse, and the mechanism is specific: automated failure learning. As the Fed's supplement puts it, an application from a 70-year-old synthetic identity with a newly established credit history is unlikely to be approved, so the system simply changes the birth year and tries again. Adversaries now A/B test underwriting policy at scale.
The supply side has industrialized alongside it. ACFE's 2026 reporting documents a Telegram tool that appeared in January 2025 offering fully automated synthetic identity creation for $250 a month, scanning for unissued Social Security numbers, planting fabricated identities into public records, and generating supporting documentation — collapsing months of manual identity aging into minutes. The same underground economy sells packaged profiles projecting credit scores above 780. This is the same trajectory visible in generative AI document fraud, where the cost of producing a convincing artifact has fallen faster than the cost of detecting one.
Where CIP Stops Being Enough
The Customer Identification Program rule requires a reasonable belief that the institution knows the true identity of each customer, and states explicitly that a bank need not establish the accuracy of every element. A well-built synthetic satisfies every element. CIP validates that identity elements exist and cohere. It does not ask whether a person corresponds to them.
The June 2025 order permitting banks to collect taxpayer identification numbers from a third party rather than from the customer removes one more point at which an applicant must personally assert a Social Security number. The underlying reasonable-belief standard is unchanged, which means the entire burden shifts onto verification quality. Institutions taking advantage of the reduced onboarding friction should be tightening the back end at the same time, alongside their broader customer due diligence and document verification controls.
The False Positive Problem Is a Fair Lending Problem
Every high-signal synthetic indicator is also the profile of a legitimate customer. No property record, no vehicle registration, short credit history, recently issued phone — that describes a young adult, a recent immigrant, a returning citizen, or a domestic violence survivor who has deliberately severed their public data trail. The Fed names this risk explicitly, warning that looking only at length of credit history could unnecessarily disadvantage immigrants and formerly incarcerated people who recently gained credit access.
The realistic operating model is a scored assessment that triggers step-up verification rather than automatic decline, because the population being step-up verified contains a large majority of legitimate thin-file customers. Automated decline should be reserved for high-confidence network matches: a shared Social Security number across identities, a shared device across unrelated names, an external account funding cards in different names.
Where Sphinx Fits
Synthetic identity detection is fundamentally a link analysis problem across systems that were not built to be joined — the onboarding platform, the core, the card portfolio, the loan book, the device telemetry, the bureau file. Sphinx's compliance agents work inside those systems directly, which means an agent can pull the cross-product view behind a flagged application, assemble the linked accounts and shared identifiers, and hand an investigator a case with the network already mapped rather than a score with no explanation attached.
The judgment call at the end is still human. Distinguishing a manufactured identity from a genuinely thin-file customer has fair lending consequences, and it is not a decision to automate away. What changes is how much evidence sits in front of the person making it.
Frequently Asked Questions
How is synthetic identity fraud different from identity theft?
Identity theft uses a real person's information, and that person eventually notices and disputes the activity. Synthetic identity fraud fabricates a person who does not exist, so no one ever complains. That absence of a victim is why synthetic losses are usually misclassified as credit losses rather than fraud, and why fraud models trained on labeled data never learn the pattern.
Why do fraud models miss synthetic identities?
The Federal Reserve reported that models built to predict traditional identity fraud fail to flag 85 to 95 percent of potential synthetic identity applicants. Traditional fraudsters move quickly because they expect to be detected; synthetics are cultivated for months or years. Velocity rules and short-window behavioral models are watching the wrong time horizon.
Does eCBSV solve the problem?
No. The Social Security Administration's electronic Consent Based SSN Verification service confirms that a Social Security number, name, and date of birth match SSA records and flags a death indicator. That is a validation floor, not a detection system — it confirms the identity tuple is real without confirming that a living person is attached to it. Adoption has also stayed well below SSA's projections since enrollment opened.
Which customers look most like synthetics but are not?
Thin-file legitimate customers. Recent immigrants, young adults, formerly incarcerated people, and consumers who have deliberately reduced their public data footprint all present short credit histories, missing public records, and recently issued contact information. The Federal Reserve warns specifically that screening on credit history length alone can unfairly disadvantage these groups, which makes this a fair lending exposure and not just a customer experience one.
Where should an institution look first if it has never checked?
The existing portfolio, not the application queue. Run cross-product sweeps for shared Social Security numbers across different names, shared devices and IP addresses across unrelated accounts, and external payment accounts remitting to multiple cards under different names. The Federal Reserve identifies payment source analysis as the sharpest in-life signal, and the highest-risk synthetics are concentrated in the super prime segment rather than in subprime.

.png)