The argument, in one page
Every business register in North America is an administrative by-product. It holds what a government needed for its own purpose - incorporation, taxation, licensing - and nothing beyond it. No aggregation layer, no vendor consolidation and no model improvement can return information that was never collected. We call that hard boundary the Collection Ceiling.
Above the ceiling sits everything a payments company or bank actually needs to make a risk decision on a small business: who really owns it, where it really trades from, what it actually sells, how much it expects to process, who it buys from, whether the person signing has authority to sign. In the United States, none of this is publicly available, and none of it exists as a general-purpose business-data source an institution can query. It exists in one place only - the memory and the filing cabinet of the business owner. We call that residual the attestation layer.
The attestation layer is not an edge case. On our field-level assessment, roughly 57% of the data points required to onboard and underwrite a North American SMB merchant are attestation-only. They are not independently retrievable at scale without the applicant's cooperation, which in practice means a human asking another human. Because they can only be asked for, they are also the fields most likely to arrive incomplete, which is why compliance and operations teams spend their days sending follow-up requests rather than reviewing risk.
The industry has spent a decade optimising the layer below the ceiling, because that is the layer vendors can sell into and benchmark. Coverage percentages, match rates, straight-through-processing rates and API latency all describe the retrievable half of the problem. There is no published benchmark anywhere for the number of information requests per onboarding, the elapsed days they add, or the applications they kill. The most expensive part of merchant onboarding is the part nobody measures, because no vendor owns it.
What this paper establishes
- The ceiling is a policy artefact, not a technical limit. The same field - beneficial ownership - is filed and publicly searchable for a federally incorporated Canadian company and structurally unavailable for a Delaware LLC. Nothing about the data makes it hard. A legislature decided.
- The United States has just lowered its own ceiling, permanently. The Corporate Transparency Act was the one serious attempt to close the largest gap in the stack. On 11 August 2026, following eighteen months of an interim exemption, Treasury issued a final rule removing the reporting requirement for all entities formed in the United States and for US persons who are beneficial owners. The single state that legislated its own register covers only foreign-formed LLCs, and holds the data in a non-public database. The manual ownership request loop is likely to remain central in the largest SMB market on earth unless policy or market infrastructure changes, and no operating model should be built on the assumption that it is temporary.
- The cost is roughly ten billion dollars a year. Our model puts the annual North American cost of the attestation layer at US$10.4bn, with a defensible range of $3.5bn to $31.7bn. Institutional labour is the smallest component. Abandoned applications are the largest.
- Perpetual KYB is structurally impossible on this base. Attestation-layer data has no refresh mechanism. A merchant's stated business model is accurate on the day it is typed and decays from that moment, with no event anywhere in the world that will tell the institution it has changed.
- Autonomous agents compress the handling cost of a request. They do not remove the request. An agent cannot retrieve a fact that does not exist in any system. It can only ask the owner faster and chase more politely.
Each of the ten sections is written to stand alone. Every section opens with a one-sentence statement of its claim and closes with a takeaway that can be lifted into a board pack without the surrounding argument. Sections 01 to 03 establish the structural case. Sections 04 to 06 quantify it. Sections 07 to 10 address what to do. The model, its assumptions and its sensitivity ranges are published in full in the method note.
A register returns what a government collected, and stops
Business data infrastructure has a hard upper bound set by what governments and institutions chose to collect, and every layer built on top of it - aggregators, orchestrators, models - redistributes that same finite pool rather than enlarging it.
There is a comfortable assumption running through the business identity market: that coverage is a function of effort. Connect to more sources, normalise more formats, resolve more entities, train better models, and the gaps close. It is an assumption that suits everyone who sells into the category, because it implies that every gap is a roadmap item.
It is also wrong, and the reason is unglamorous. A company register is not a description of a business. It is the residue of a transaction between a business and a state. Delaware asks for the information Delaware needs to constitute a legal person and serve process on it. The IRS asks for the information the IRS needs to assign a tax identifier. A state licensing board asks for the information it needs to grant a licence. None of these bodies is in the business of describing a company to a third party, and none of them collects a field simply because someone downstream might one day find it useful.
This produces a structural boundary that no amount of downstream engineering can cross. If a fact was never collected by anyone, it is not in the pool. It cannot be aggregated, because there is nothing to aggregate. It cannot be inferred reliably, because inference on a firm with four employees and no web presence is guesswork with a confidence score attached. It cannot be modelled into existence. The only remaining route to that fact is the person who holds it.
Two ceilings, not one
The boundary is actually two boundaries stacked, and conflating them causes bad procurement decisions.
The collection ceiling is set by what was gathered at all. If no filing anywhere captures the percentage of an LLC held by each member, that number exists in no system.
The distribution ceiling sits below it, and is set by what is released to whom. A fact can be collected and still be unavailable: held in a non-public register accessible only to law enforcement, released only in paper form, priced beyond reach, published without a machine-readable interface, or updated on an annual cycle that makes it stale for operational use. New York's beneficial ownership regime is the cleanest example in North America: the state legislated a register, and then specified that the data sits in a database available only to government agencies. Collected, and still invisible to the bank.
For an operator the two ceilings feel identical - the field is missing either way. For a policymaker they are entirely different problems with entirely different remedies, and a paper that blurs them produces recommendations that cannot be acted on. We keep them separate throughout.
Why the ceiling is invisible from inside a vendor
Every metric the business verification market reports is a ratio with the retrievable universe as its denominator. Coverage is measured as jurisdictions connected, not as fields required. Match rate is measured against records that exist, not against the decision the institution has to make. Straight-through-processing rate counts cases that cleared automatically, which by construction excludes every case that was routed to a human precisely because the data was not there.
The result is a market that reports excellent and improving numbers about a shrinking share of the actual problem. A provider can move from 82% to 94% coverage of state filing records and change the compliance team's working day almost not at all, because the working day is not spent retrieving filing records. It is spent chasing the thirty-one fields that were never in a filing record to begin with.
The test that separates the two layers
A field is below the ceiling if a stranger with the right credentials and enough money could obtain it without the company's cooperation. A field is above the ceiling if the only path to it runs through the company itself. That test is deliberately indifferent to how important the field is, how sensitive it is, or how obvious the answer seems. It asks one thing: does an independent record exist?
Apply it honestly and the results are uncomfortable. A merchant's legal name passes. Its operating address usually fails, because what the register holds is a registered office or a registered agent's address, which is frequently neither where the business trades nor where its goods are held. Its ownership fails almost everywhere in the United States. Its business model fails in almost every case: licensing, tax and sectoral regimes capture fragments of business purpose or activity, but no register captures what a company actually sells in a form that is structured, current and comparable across jurisdictions.
The business data industry has been solving a distribution problem while describing it as a coverage problem. Below the collection ceiling, consolidation genuinely helps and the market has made real progress. Above it, vendor capability converges: no provider can create a standardised authoritative source for a fact that was never collected or never released. The competitive question is no longer who retrieves the most, but who handles the residual best.
Sixty-four registers, no federal one, and a reversal on ownership
Company formation in the United States is a state function spread across more than sixty subnational registers with no federal equivalent, and the one federal attempt to capture beneficial ownership has now been reversed, leaving US institutions to source ownership data entirely from the business itself.
Company formation in the United States is a state function. Fifty states and the District of Columbia each operate their own register, with their own fields, formats, fee schedules, update cadences and access methods. Canada adds ten provincial and three territorial registries, plus a federal register under the Canada Business Corporations Act. Throughout this paper, the sixty-four filing jurisdictions are the fifty US states, the District of Columbia, the ten Canadian provinces and the three Canadian territories. US territories and the Canadian federal register sit outside that count and are discussed separately where relevant. An institution onboarding North American merchants is therefore reconciling sixty-four distinct subnational data regimes, none of which was designed to interoperate with the others and none of which was designed with a bank in mind.
Fragmentation is the familiar complaint, and it is real, but it is the lesser problem. Fragmentation is expensive; it is not a ceiling. Enough engineering and enough spend can normalise sixty formats into one schema, and several providers have done exactly that to a high standard. What engineering cannot do is add fields to sixty registers that never held them. Even a perfectly unified North American view of every state and provincial filing would return: legal name, entity type, jurisdiction, formation date, standing, filing number, registered agent, registered address, and in some states a partial and often stale list of officers. That is the ceiling. It is a thin dossier on which to make a payments risk decision about a business that intends to process seven figures a year.
The Corporate Transparency Act, and its reversal
The Corporate Transparency Act was the one federal attempt to raise the ceiling on the single most consequential missing field. From January 2024 it required tens of millions of US entities to file beneficial ownership information with FinCEN: name, date of birth, address and identification document for every individual holding 25% or more or exercising substantial control.
It did not survive. Following the Treasury announcement of March 2025 and the interim final rule published on 26 March 2025, the definition of a reporting company was narrowed to entities formed under the law of a foreign country and registered to do business in a US state or tribal jurisdiction. All entities created in the United States, and their beneficial owners, were exempted. Commentators put the practical effect at the removal of roughly 99.8% of originally covered entities from scope. The exemption is now final. On 11 August 2026, Treasury issued a final rule solidifying the position: entities formed in the United States and US persons who are beneficial owners are permanently outside the reporting requirement. The rule additionally exempts foreign companies from reporting US-person company applicants, and removes the obligation on US persons to update or correct the information behind a FinCEN identifier. Treasury grounded the decision in the statute's own instruction to minimise the burden that collection places on legitimate businesses, describing the outcome as a balancing test between generating useful information and limiting compliance cost.
We take no position on whether that was the right policy. The operational consequence is not in dispute. For a US-formed small business - which is to say very nearly every small business a North American payments company will ever onboard - there is no authoritative source of beneficial ownership. There is no register to query, no file to fetch, no vendor with privileged access. There is a form, sent by a compliance analyst, asking the owner to declare it. And then, in a material share of cases, a second form.
The burden that was not removed
Treasury's reasoning is the clearest possible statement of the argument this paper makes, arrived at from the opposite direction. The department weighed the cost of collection against its usefulness and concluded that collection imposed too much burden on legitimate business. That trade-off is real. But the obligation on financial institutions to identify beneficial owners was not withdrawn alongside it, and neither was the small business owner's obligation to answer.
What the final rule removes is one structured filing, made once, at formation. What it leaves in place is the same disclosure, made repeatedly, in a different format, to every bank, acquirer, lender and marketplace the business ever approaches, for the life of the business, with no persistence and no reuse. The burden was not eliminated. It was moved off the government's ledger and onto a ledger nobody keeps.
Critics of the rule have framed it as a law-enforcement and transparency question, and it is one. It is also an operating cost question, and that dimension has gone almost entirely unexamined. This paper puts the North American cost of the resulting request loop at $10.4bn a year.
The state-level backstop that closed before it opened
New York is the only state-level regime within this paper's scope that requires beneficial-owner disclosure from a defined class of LLCs, modelled on the federal act. The LLC Transparency Act took effect on 1 January 2026. Two features of how it landed matter more than the fact of its existence.
First, its key definitions are tied to the federal act. The Governor's December 2025 veto of the decoupling amendment meant those definitions inherited the federal exemptions, so the New York register now applies only to LLCs formed outside the United States and authorised to do business in the state. US-formed LLCs, including New York LLCs, are out of scope.
Second, and more instructive for anyone designing a remedy, the disclosures are held confidentially rather than on the public record, and current Department of State guidance provides an access route for law enforcement rather than for supervised financial institutions. Even for the narrow population in scope, an institution conducting customer due diligence has no route to it. This is the distribution ceiling in its purest form: the state collected the data and closed it to the institutions carrying the obligation it was meant to serve.
Canada as the controlled experiment
The most useful thing about North America as a study region is that it contains its own counterfactual. Since 22 January 2024, corporations governed by the Canada Business Corporations Act have been required to file information on individuals with significant control with Corporations Canada, and a portion of that data is publicly searchable: name, jurisdiction of residence, description of control, and the date control began, with residential addresses withheld. Quebec goes further still, requiring ultimate-beneficiary disclosure on a publicly accessible enterprise register, extending to foreign entities that are required to register in the province.
The same field, for the same kind of business, at the same moment in history: filed and queryable in Ottawa and Quebec City, and structurally unobtainable in Dover, Albany and Sacramento. That comparison is the whole argument in miniature. Nothing about beneficial ownership makes it technically hard to collect. The ceiling is where a legislature put it.
Canada is not a solved market. Provincial coverage is uneven, several provinces require only internal registers that are never filed, and the federal register's population depends on annual return cycles. But the direction of travel is opposite to the United States, and any institution operating across the border now runs two materially different KYB realities under one policy.
Beneficial ownership availability, North America, mid-2026
| Jurisdiction class | Collected? | Available to an FI? | Practical position |
|---|---|---|---|
| US federal (FinCEN, domestic entities) | No | No | Permanently exempted by final rule, August 2026 |
| US federal (FinCEN, foreign-formed entities) | Partial | Conditional | Certain foreign reporting companies only, and only non-US-person ownership. Not public; institutional access is consent-based and limited |
| US states, 49 of 50 | No | No | No comprehensive beneficial-ownership register accessible to financial institutions |
| New York | Partial | No | Foreign-formed LLCs only; non-public database |
| Canada federal (CBCA) | Yes | Partial | ISC filings, portion publicly searchable |
| Quebec | Yes | Yes | Public enterprise register, incl. foreign entities |
| Other provinces | Mixed | Mostly no | Internal registers, not filed centrally |
An institution onboarding US SMBs in 2026 has less structured ownership data available to it than it expected to have in 2024, while carrying the same obligation to know. That gap did not close through better tooling. It was reopened by policy, and it will stay open until policy changes, which means every operating model built between now and then has to assume the request loop is permanent.
Fifty-four fields, and thirty-one of them are locked in someone's filing cabinet
Mapping the data a North American payment service provider actually requires to onboard and underwrite an SMB merchant against every source that could supply it shows that 57% of required fields have no independent record anywhere and can only be obtained by asking the business.
The argument so far is structural. This section makes it countable. We took the composite requirement set that a North American acquirer, payment service provider or business bank applies to an SMB merchant - identity, ownership, operating reality, financial profile, counterparty exposure and screening - and reduced it to fifty-four discrete data points. We then classified each one by the highest-confidence source that could supply it without the merchant's involvement.
Four tiers, applied consistently:
- Tier 1 - Registry-derivable. Held in a state, provincial or federal company register and retrievable, in principle, by anyone.
- Tier 2 - Other authoritative record. Held by a tax authority, securities regulator, court, licensing body or sanctions list. Often verification-only: the source will confirm a value you supply, but will not hand you the value.
- Tier 3 - Commercially inferable. Available from credit bureaux, firmographic providers, trade payment panels or the open web, with a confidence score attached. Confidence collapses at the small end, which is precisely the segment under discussion.
- Tier 4 - Attestation layer. No generally available, independently retrievable authoritative source exists at commercial scale. The field can be obtained only with the applicant's cooperation.
The tier assignment is made against a US-formed private company, because that is the modal merchant. Where a Canadian federal or Quebec entity would sit in a different tier, we note it. That divergence is itself one of the findings.
Scope. This is a composite requirement set for a higher-friction North American SMB use case: merchant acquiring and payment acceptance, business banking, and small business lending. It is not a universal KYB requirement set. A marketplace onboarding a service provider, a payroll platform, or a low-risk deposit-only relationship will require materially less. Every figure in this paper, including the 57%, describes that composite population and should not be read across to lighter use cases.
Two kinds of attestation-layer field. The attestation layer contains both customer declarations, which are statements the applicant makes about itself such as ownership percentages, expected volume or business model, and customer-supplied evidence, which are artefacts the applicant produces such as bank statements, processor statements, financial statements and supplier invoices. They differ in how they are verified and in who bears the effort of producing them. They behave identically for the purposes of this paper, because both can only be obtained by asking, and the economic model in Section 05 covers both.
The onboarding field map
| Group | Fields | T1 | T2 | T3 | T4 | Representative attestation-only fields in this group |
|---|---|---|---|---|---|---|
| A. Legal identity | 9 | 8 | 1 | 0 | 0 | - the one group the ceiling does not bite |
| B. Tax & federal identity | 4 | 0 | 4 | 0 | 0 | verification-only; values must still be supplied first |
| C. Ownership & control | 9 | 1 | 0 | 0 | 8 | ownership percentages, ownership chain, control persons, signatory authority, affiliated entities, nominee arrangements |
| D. Operating reality | 9 | 0 | 0 | 1 | 8 | trading address, products actually sold, business model, fulfilment and delivery, refund policy, acquisition channels, seasonality |
| E. Financial profile | 10 | 0 | 1 | 0 | 9 | expected volume, average ticket, processor statements, chargeback history, bank statements, financial statements |
| F. Counterparty & supply chain | 5 | 0 | 0 | 2 | 3 | key suppliers, inventory sourcing evidence, fulfilment partners |
| G. Risk & screening | 8 | 0 | 4 | 1 | 3 | declared industry classification, source of funds, owner screening (dependent on Group C) |
| Total | 54 | 9 | 10 | 4 | 31 | 57% of the requirement |
Attestation share by field group
Tier assignment is an expert structural assessment, not a survey of every jurisdiction. It reflects the highest-confidence source reasonably obtainable at commercial scale for a privately held US SMB in mid-2026. Where a field is available in some jurisdictions but not most, it is assigned to the tier that governs the majority of onboarding events. A full field-level matrix carrying per-jurisdiction notes for all sixty-four North American filing jurisdictions would be a more precise instrument than the one presented here, and does not currently exist in the public domain. Building it is identified as follow-up research in the method note.
Three things the matrix makes visible that a coverage percentage hides
Verification is not retrieval
Tier 2 flatters the numbers if read carelessly. A tax authority will tell you whether a name and an EIN belong together. It will not tell you the EIN. Bank account ownership can be confirmed against an account you already have. Sanctions screening operates on a name you have already been given. Almost every Tier 2 field is a check on a value the merchant supplied, which means the merchant still had to supply it, which means the request still had to be sent. Counted honestly, the number of fields an institution can populate with zero merchant input is nine.
The dependency chain runs the wrong way
Screening beneficial owners against sanctions and politically exposed person lists is a Tier 2 capability sitting on a Tier 4 input. The screening technology is mature, cheap and instant. It is useless until a human has typed in the names. In the United States that means the entire sanctions posture on privately held SMB merchants rests on a self-declaration that no register can corroborate. This is the single most consequential structural fact in the paper, and it is a policy outcome rather than a technology gap.
The gap widens as the business gets smaller
Tier 3 inference is the layer that is supposed to cover the middle ground, and it works reasonably on a company with audited accounts, a trade payment history and a digital footprint. The modal North American small business has none of these. Roughly four in five US small businesses are nonemployer businesses. A firm formed last quarter has no trade panel history, no bureau file worth the name, and a website that went live three weeks ago. New and thin-file businesses are therefore the ones on which commercial inference performs worst, which pushes them further above the ceiling rather than below it.
Procurement conversations in this category are conducted almost entirely about nine fields. The decision that actually determines onboarding performance is how the institution handles the other thirty-one, and that decision is usually made implicitly, by default, in a ticketing queue nobody owns.
What happens above the ceiling: the request loop, and why it compounds
Because attestation-layer fields can only be requested, onboarding becomes an asynchronous negotiation between an institution and a business owner in which each round trip costs days, some round trips are triggered by omissions no institution could have prevented while others are self-inflicted, and the number of round trips - not the review itself - sets the total elapsed time.
The operating pattern is consistent across every institution we have observed, from twelve-person fintechs to global acquirers. A merchant applies. The institution retrieves what it can below the ceiling, presents a form for everything above it, and the merchant fills it in with whatever is to hand. An analyst reviews the submission, finds gaps - missing owner, unclear business model, statements from the wrong months, an address that does not match the registry, a document that is legible but not the right document - and sends a request. The merchant responds, usually not immediately. The analyst reviews again. Frequently a second gap surfaces, sometimes revealed by the answer to the first.
Some of this is irreducible. Any process whose inputs live outside the system of record must ask for them, and the institution cannot validate a field it has not received or know which fields will be wrong until they arrive. But not every round trip is irreducible. Incomplete requirement design, weak prefill, fragmented case management and context loss between rounds generate avoidable requests on top of the necessary ones, and Section 09 sets out what to do about those.
Anatomy of one cycle
A single request cycle contains six distinct costs, only one of which is usually measured.
- Detection. An analyst identifies the gap. This is the only step that looks like compliance work.
- Composition. The request is drafted, often into a template that does not carry the case context, and dispatched.
- Latency. The clock runs while the merchant is doing something else. This dominates elapsed time and is entirely outside the institution's control.
- Retrieval by the merchant. The owner locates a statement, asks an accountant, photographs a document, or asks a co-owner for a date of birth. This is unpaid work performed by the customer.
- Re-review. The analyst reconstructs the case from memory or from notes, because in most stacks the follow-up arrives as a new ticket with no thread back to the original assessment.
- Recurrence. A material share of responses generate a further gap, and the loop restarts.
One request cycle
- 1DetectionAnalyst identifies the gap. The only step that looks like compliance work.
- 2CompositionRequest drafted into a template that does not carry case context, and dispatched.
- 3LatencyThe clock runs while the merchant is doing something else. Dominates elapsed time. Entirely outside institutional control.
- 4Retrieval by the merchantOwner finds a statement, asks an accountant, chases a co-owner. Unpaid work performed by the customer.
- 5Re-reviewAnalyst reconstructs the case. The follow-up usually arrives as a new ticket with no thread back to the original assessment.
- 6RecurrenceA material share of responses generate a further gap, and the loop restarts from step one.
Step five deserves particular attention because it is the most fixable and the least discussed. In several large institutions each additional information request opens a fresh case object with no context inherited from the previous one. An analyst who has already spent twenty minutes understanding a merchant's corporate structure spends much of it again on the next touch. The cost of a request is therefore not linear in the number of requests; it is superlinear, because context is discarded between rounds.
Observed case: a global payment service provider
The following is drawn from internal operating data shared with Detected under confidentiality by a global payment service provider, and is presented in abstracted form. Absolute volumes, product names, internal system names and segment identifiers are withheld. Ratios and durations are reproduced as reported. Scope and definitions, as supplied: North American SMB applications across more than one acceptance product, first half 2026; durations are reported as averages rather than medians or percentiles, and as calendar days; the underwriting figure covers applications entering the full review band and does not separate out integration-dependent cases; the decline share counts applications refused on fraud or risk grounds and excludes applications that lapsed, were withdrawn, or were closed as technically incomplete. This is a single non-representative illustrative case. It is presented because the pattern matches what we observe across institutions, not because one operator can stand for a market.
The operator runs SMB merchant onboarding across more than one acceptance product, splitting applicants into review bands by expected processing volume: a lighter risk-scored review for the lower band, and a full underwriting and compliance review above it. The published customer-facing experience is a self-serve account creation flow that takes about ten minutes.
The gap between that ten minutes and the actual time to first transaction is the entire subject of this paper.
Reported stage timings, SMB onboarding
| Stage | Reported duration | What governs it |
|---|---|---|
| Account creation | ~10 minutes | Form design. Fully self-serve. |
| Identity programme review, clean documents | 24 business hours | Analyst throughput |
| Identity programme review, missing or unclean documents | up to 1 week | Number of merchant touchpoints |
| Complex acceptance integrations | up to 3 months | Document exchange and technical dependency |
| Underwriting and compliance review | ~70 days average | Document exchange and queue |
| Signup to first transaction | ~70 days | Attestation-layer exchange |
Three details in this dataset carry more weight than the headline.
The stated cause of the difference between one day and one week is touchpoints, not workload. The operator's own characterisation of the identity review stage distinguishes clean submissions, resolved within a business day, from submissions requiring multiple merchant touchpoints, which take up to a week. Same analysts, same queue, same regulation. The variable is how many times the merchant had to be asked.
Context loss between requests is named internally as a bottleneck. Each additional document requirement generates a new internal ticket, and context is not carried between tickets. The operator identifies this explicitly as adding days to resolution. This is step five above, observed in the wild at scale, in an organisation with world-class engineering resources. It persists because it is nobody's KPI.
The document list above the ceiling is long and almost entirely attestation-only. The full review band requires two years of audited or company-prepared financial statements or corporate tax returns; interim year-to-date financials; three months of bank statements; processor statements; organisational chart and corporate structure; inventory receipts or supplier information; products, services and pricing; professional licences; and key vendors and suppliers. Not one of these is retrievable from any register. Every one is a potential request cycle.
The ratio that should trouble the industry
In the same reporting period, the share of SMB applications declined on fraud or risk signals was under 5%.
Read that against a seventy-day average. A process consuming roughly ten weeks of elapsed time, on both sides of the relationship, resolves to a decline in fewer than one case in twenty. The remaining applications were ultimately approved. We cannot observe what would have happened to them under a different process, but they were not under active assessment for most of those seventy days; they were waiting for a document, or waiting for someone to notice a document had arrived.
Reported churn during the process was under 4% at this operator - a managed segment with named account teams actively shepherding merchants through. That is a managed best case rather than a typical one. We are not able to quantify attrition in unmanaged self-serve SMB onboarding, where nobody chases, and we have seen no published benchmark for it. We would expect it to be materially higher, and the abandonment range in Section 05 is set wide for that reason.
Onboarding duration is a function of request count, and request count is a function of how much of the requirement sits above the collection ceiling. Institutions optimising analyst productivity are optimising a step that consumes a small fraction of elapsed time. The addressable variable is the number of round trips, and the cheapest reduction available to most institutions today is not a new data source but carrying case context between them.
Sizing the attestation layer
On a transparent model with published assumptions, the annual cost of the attestation layer to the North American economy is approximately US$10.4bn, of which institutional compliance labour - the only component anyone currently budgets for - is around 13%.
There is no survey to cite here, because none has been conducted. What follows is a model, not a measurement, and we present it as such: every input is stated, every input is arguable, and the ranges are wide enough to survive most disagreements about them. Readers who dislike our numbers should replace them and re-run the arithmetic. The structure of the result - that abandonment and deferred activation together outweigh institutional labour by roughly five and a half to one - is robust across the whole plausible parameter space, and that structural finding matters more than the headline.
Step one: how many onboarding events
An onboarding event is one SMB being taken on by one regulated financial or payments provider, requiring a KYB and customer due diligence assessment. A business opening a bank account and separately taking on card acceptance generates two events.
The United States records roughly 5.4 million business applications a year. Most never become operating businesses with a payments relationship, so we anchor instead on high-propensity applications, which run at approximately 1.9 million annually, and add a Canadian equivalent of around 0.25 million. Each new operating business establishes an average of 1.9 regulated relationships in its first year - typically a business bank account and one acceptance or lending product. That gives 4.1 million first-year events.
The installed base of roughly 38 million North American small businesses adds a second layer. We assume 12% add or switch a regulated provider in any given year, at an average of 1.12 new relationships each, giving 5.1 million switching and expansion events.
Total: 9.2 million SMB onboarding events per year across North America, with a range of 6.8 to 12.4 million. We deliberately exclude periodic refresh and remediation cycles, which run the same request loop and would materially increase the total. They are discussed separately in Section 06.
Model inputs
| Input | Low | Central | High | Basis |
|---|---|---|---|---|
| Annual SMB onboarding events, North America | 6.8m | 9.2m | 12.4m | Census BFS high-propensity applications; SBA installed base |
| Follow-up requests per onboarding | 1.7 | 2.3 | 3.6 | Practitioner observation across acquiring, banking and lending |
| Institution minutes, initial attestation review | 45 | 55 | 70 | Analyst review of merchant-supplied material |
| Institution minutes per follow-up cycle | 32 | 38 | 48 | Detection, composition, tracking, context reconstruction, re-review |
| Loaded institutional cost per hour | $55 | $62 | $72 | Fully loaded compliance operations, US and Canada |
| Merchant minutes, initial assembly | 80 | 95 | 120 | Locating and producing documents and declarations |
| Merchant minutes per follow-up cycle | 55 | 71 | 95 | Including third-party dependencies such as accountants and co-owners |
| Merchant opportunity cost per hour | $38 | $42 | $52 | Blended owner and administrator time |
| Elapsed days added per follow-up cycle | 3.4 | 4.7 | 7.2 | Dominated by merchant response latency, not analyst throughput |
| Abandonment attributable to information requests | 6% | 11% | 18% | Between a managed-segment floor under 4% and unmanaged digital onboarding attrition |
| Provider lifetime revenue per SMB relationship | $3,400 | $4,600 | $5,400 | Net revenue over average tenure, discounted |
| Merchant daily gross profit at risk during delay | $260 | $310 | $420 | SMB in the process of adding an acceptance or banking relationship |
| Perishability of delayed activity | 8% | 10% | 10% | Share of delayed trading permanently lost rather than deferred |
Step two: the four cost pools
Annual cost of the attestation layer, North America
| Cost pool | Annual cost | Share | Who bears it |
|---|---|---|---|
| Institutional compliance labour | $1.35bn | 13% | Banks, acquirers, PSPs, lenders |
| Merchant labour | $1.66bn | 16% | Small business owners, unpaid |
| Deferred activation | $2.75bn | 26% | Merchants, and providers through delayed revenue |
| Attributable abandonment | $4.65bn | 45% | Providers, as forgone lifetime revenue |
| Total | $10.41bn | 100% |
Formulas. Institutional labour = events × [initial review minutes + (requests × follow-up minutes)] ÷ 60 × loaded hourly cost. Merchant labour = events × [initial assembly minutes + (requests × follow-up minutes)] ÷ 60 × merchant hourly cost. Deferred activation = events × (1 − abandonment rate) × requests × days per request × daily gross profit at risk × perishability. Attributable abandonment = events × abandonment rate × provider lifetime revenue.
On double counting. Deferred activation is measured as merchant gross profit forgone during delay, and applies only to the onboardings that complete. Attributable abandonment is measured as provider lifetime revenue forgone, and applies only to the onboardings that do not. The two pools are therefore disjoint by construction and cannot double count the same relationship. They are summed because they fall on different parties, and neither is captured in the labour pools.
The cost is not where the budget is
The only pool with a budget line is the smallest one. The industry manages 13% of the problem and reports on it monthly.
Step three: sensitivity
Setting every input simultaneously to the low column of Table 5.1 produces $3.5bn. Setting every input simultaneously to the high column produces $31.7bn. These are the arithmetic extremes of the stated inputs rather than a percentile interval, and no correlation between inputs is modelled, so the true distribution is almost certainly narrower than the band. The width is honest: nobody has measured request counts at industry scale, and that single parameter drives most of the variance. What does not change across the range is the ordering. Abandonment is the largest pool in every scenario. Institutional labour is the smallest in every scenario. An industry that has organised its cost management around the smallest pool, because it is the only one that appears on a budget line, is optimising against a total nearly eight times its size.
Sensitivity to requests per onboarding
What the model deliberately excludes
- Periodic refresh and remediation. Every KYC refresh cycle re-runs the same loop on the same attestation-layer fields. We have not modelled it, because refresh cadence varies by risk band and institution and we have no basis for an industry figure.
- Financial crime losses attributable to unverifiable attestations. Unknowable without incident-level data, and we decline to estimate it.
- The cost of the businesses that never applied. Discussed qualitatively in Section 06; not quantified, because the counterfactual is unobservable.
- Enterprise and mid-market onboarding. The same structure applies with more resources on both sides. This model covers SMB only.
Several of these exclusions could increase the estimate. We have not quantified them because they may overlap with the modelled pools or cannot be isolated reliably, and we would rather leave them out than add figures we cannot defend. On that basis we regard $10.4bn as a conservative reading of a problem nobody has previously bothered to size.
The attestation layer is a ten-billion-dollar annual cost centre in North America that appears on no budget, in no vendor benchmark and in no regulatory impact assessment. Four fifths of it lands outside the compliance function that causes it - on merchants, on commercial teams, and on revenue that never arrives.
Five costs, only one of which is on a budget line
Beyond direct expense, the attestation layer selects the wrong customers, guarantees that risk data decays unobserved, and excludes the businesses with the thinnest registry footprint - which are systematically the newest, smallest and least established.
One. Operational drag that scales with growth, not with risk
Request volume is a function of merchant count, not of merchant risk. Doubling merchant acquisition doubles the number of request cycles regardless of whether the incremental merchants are riskier. This is why compliance headcount grows in step with commercial success and why the relationship feels, to every commercial leader who has lived it, like a tax on growth rather than a control. It is a tax on growth. The control is the review; the tax is the retrieval.
It also concentrates. Because most request cycles resolve into approvals, the analyst hours are disproportionately spent on merchants who were never going to be declined. The institution's scarcest expert resource is allocated by document completeness rather than by risk.
Two. Deferred and forgone revenue on both sides
For the provider, an application in the request loop is inventory. It has been paid for through acquisition spend, it consumes servicing cost, and it generates nothing. Extending average time to activation by ten days across a merchant book is a working capital event, not a customer experience event.
For the merchant, the delay lands at the worst possible moment. A business seeking payment acceptance is usually seeking it because it has demand it cannot currently serve. Ten weeks between signing up and taking a first payment is, for a seasonal business, an entire season.
Three. Adverse selection, which is the finding that should worry regulators most
A multi-round information request process may reward persistence more than it rewards integrity.
A legitimate owner-operator running a busy small business responds to a fourth document request with irritation and, often, with abandonment. They have alternatives, they are busy, and the request feels like a judgement. A professional bad actor responds to the fourth request promptly and completely, because completing onboarding is their objective and because fabricated documents are cheap to produce. The material required - a lease, an invoice, a statement, an org chart - is precisely the material that is easiest to falsify and hardest to verify at scale, because there is no independent record to check it against. That is what being above the ceiling means.
Friction, in this specific setting, should therefore not be assumed to be a proxy for control. It may disadvantage time-constrained legitimate applicants without improving risk selection. We do not claim the effect is dominant, and we cannot claim it is established, because we have not measured it. We state it as a hypothesis running counter to the intuition everyone holds, and we found no published study measuring completion rates split by eventual risk outcome. That study, proposed in the method note, is the test that would settle it.
Four. Risk decay, and why perpetual KYB cannot be built on this base
The industry has largely accepted that periodic review is inferior to continuous monitoring, and vendors have responded with event-driven products: registry change alerts, sanctions list updates, adverse media triggers, litigation and filing events. All of these monitor below the ceiling.
Above it, monitoring is structurally incomplete. A merchant declares an expected annual volume, an average ticket, a product range and a fulfilment model on the day of application. Each of those is accurate for as long as it is accurate, and no authoritative event exists that will tell the institution when it stops being true. A merchant that pivots from selling homeware to selling supplements, changes fulfilment from held stock to drop-shipping, or begins accepting pre-orders with six-month delivery windows has materially changed its risk profile without triggering any filing, listing or registry event. Partial signals may surface from transaction patterns, website changes, shipment data, complaints or adverse media, and mature programmes use them. They are lagging, non-universal, and rarely sufficient on their own to support the determination the institution is required to make, which means the practical detection mechanism remains transactional: the institution finds out from the chargebacks. The August 2026 final rule extends the same problem to the narrow population still in scope: by removing the obligation on US persons to update or correct the information behind a FinCEN identifier, it guarantees that even the data that continues to be collected will decay from the day it is filed.
This is a structural limit on perpetual KYB, and it is under-acknowledged in a market that sells continuous monitoring as though it were complete. Continuous monitoring cannot reliably refresh a customer-supplied fact without either a new independent source or a renewed confirmation from the customer. Monitoring the retrievable layer well is a genuine improvement over reviewing it annually. It says much less about the thirty-one fields that actually describe what the business does.
What continuous monitoring actually covers
- Registry changes
- Sanctions and PEP
- Adverse media
- Litigation and filings
- Business model
- Products sold
- Expected volume
- Fulfilment model
- Ownership changes
- Supplier relationships
Above the ceiling, partial and lagging signals may exist, but no authoritative event source does. In practice the detection mechanism is transactional.
Five. Financial exclusion, concentrated exactly where policy says it should not be
The burden of the attestation layer is not evenly distributed. It falls hardest on businesses with the thinnest independent record, and the thinness of a business's record correlates with characteristics that have nothing to do with risk.
- Age. A firm formed six months ago has no trade payment history, no bureau depth and minimal web presence. Every Tier 3 inference fails, and everything falls back to attestation.
- Size. Around four in five US small businesses are nonemployer businesses, with no payroll beyond the owner, on SBA Office of Advocacy figures for 2025. There is no finance function to produce interim financials on request, and no administrator to chase a supplier for a letter.
- Structure. Sole proprietorships and partnerships frequently have no registry record at all, so even the nine Tier 1 fields are unavailable and the entire dossier is attestation-derived.
- Documentary fluency. Producing an organisational chart, a year-to-date balance sheet and a supplier schedule on request is a skill unevenly distributed across founders, and correlated with prior experience of formal institutions rather than with the quality of the business.
The consequence is that the merchants most likely to abandon are the newest, smallest and least institutionally fluent. They are also, on SBA Office of Advocacy figures, the population responsible for a substantial share of net new job creation in the United States. An onboarding process that filters this population out through documentary attrition is producing an exclusion outcome that no regulator intended and no institution measures. Nobody is doing anything wrong. That is precisely why it persists.
The attestation layer produces four consequences that no institution currently instruments: it allocates expert time by document completeness rather than risk, it selects mildly in favour of the well-resourced bad actor, it makes continuous monitoring structurally incomplete, and it excludes the newest and smallest businesses through attrition rather than decision. Each of these is measurable today by any institution willing to split its funnel by eventual risk outcome. Almost none do.
Why the largest cost in onboarding has no owner and no metric
The attestation layer is unmeasured because it falls between every function that could measure it, and unbenchmarked because no vendor can sell against it, which means the only actor with an incentive to quantify it is the institution paying for it.
Problems get measured when someone can be paid for solving them. That is the whole explanation, and it is worth spelling out because it also indicates where the measurement will come from.
Vendors benchmark what they can win on
Every published metric in the business verification category describes retrieval: jurisdictions covered, match rate, straight-through-processing rate, API latency, data freshness. These are real capabilities and providers compete hard on them. But a metric only becomes an industry benchmark when at least two vendors want to be compared on it. No provider can differentiate on requests per onboarding, because no provider controls the field that triggers the request. The metric that governs the cost is the one metric nobody in the supply chain benefits from publishing.
The cost is split across four functions, and owned by none
Analyst hours sit in compliance. Elapsed time to activation sits in operations. Abandonment sits in growth or commercial. Merchant effort sits nowhere at all, because it is borne by the customer. Each function sees a fragment: compliance sees a manageable queue, operations sees a cycle time it attributes to compliance, growth sees a conversion rate it attributes to product. The aggregate is visible only from a vantage point that most organisations do not have, and the diagnosis - that the true driver is what the state did not collect - sits outside the remit of all four.
Straight-through-processing rate actively conceals it
STP rate is the industry's headline efficiency measure, and it is defined as the proportion of cases completing without human intervention. Cases requiring information requests are, by definition, excluded from the numerator. An institution can raise STP from 61% to 74% and see no change whatever in average time to activation, because the cases that set the average are the ones STP excludes. The metric improves as the problem is pushed further out of view.
Regulatory reporting is about outcomes, not effort
Supervisors ask whether the institution knows its customer and whether its programme is effective. They do not ask how many times it had to write to the customer to find out, and we found no standard North American supervisory disclosure framework that captures it. The effort is therefore invisible to the actor best placed to change the underlying policy.
Nobody assigns a cost to unpaid customer time
The single largest pool of hours in this system - nearly forty million a year - is worked by small business owners for free. It appears in no P&L on either side. It is a genuine economic cost, borne by the part of the economy least able to absorb it, and it is structurally invisible because there is no transaction attached to it.
Who sees what
| Function | Fragment it holds | What it concludes |
|---|---|---|
| Compliance | Analyst hours | A manageable queue |
| Operations | Elapsed time | A cycle time it attributes to compliance |
| Growth | Abandonment | A conversion rate it attributes to product |
| The merchant | Unpaid hours | An ordeal, and no invoice |
| Nobody | The aggregate | $10.4bn |
Four metrics that would make it visible
| Metric | Definition | Why it matters |
|---|---|---|
| Requests per onboarding | Mean information requests sent after initial submission | The single strongest predictor of elapsed time to activation |
| Attestation share | Proportion of required fields with no independent source | Sets the ceiling on any automation programme before it starts |
| Request-attributable abandonment | Applications abandoned within 14 days of a request, as a share of requests | Converts a compliance process into a revenue number |
| Completion split by risk outcome | Completion rate of applicants later flagged, versus never flagged | Tests directly whether friction is selecting for integrity or against it |
Any institution can begin measuring this quarter, using data already in its case management system. The first institution in a given market to publish requests-per-onboarding will define the benchmark, and will discover its own number is worse than it assumed.
Four remedies that will not raise the ceiling
More data sources, more vendors, better models and autonomous agents all operate below the collection ceiling or on the handling cost of requests; none of them can produce a fact that no party ever recorded.
More data sources
The reflexive institutional response to a coverage gap is to add a provider. Where the gap is genuinely a distribution problem - a jurisdiction not connected, a format not parsed - this works, and it is the right answer. Where the field is above the ceiling, adding a fifth provider to four existing ones adds cost, integration surface and reconciliation work while returning the same null. The diagnostic question before any procurement decision in this category is simply: is this field held by anybody at all? If the answer is no, no contract will change it.
Vendor consolidation and orchestration
Orchestration is genuinely valuable. Routing to the best source per jurisdiction, falling back gracefully, reconciling conflicting records and maintaining one schema across sixty regimes is hard, and doing it well materially reduces cost and latency below the ceiling. But orchestration is a distribution technology by definition. Consolidating five providers into one improves the economics and the engineering of retrieving the nine Tier 1 fields. The thirty-one Tier 4 fields are unaffected, because they were never in any of the five.
Institutions consistently overestimate what consolidation will do to cycle time for this reason. They measure the improvement in the retrieval step, which is real, and are then surprised that end-to-end time barely moves, because the retrieval step was never the constraint.
Better inference
Inference from digital footprint, transaction patterns and network signals is improving quickly and will continue to. Two limits apply. First, it degrades exactly where the volume is: the newest and smallest businesses have the thinnest signal, and they generate the most onboarding events. Second, an inferred value is not an attested value, and for a regulated determination the institution needs a record of what the customer stated, not a probability. Inference can prioritise a queue, flag an inconsistency and pre-fill a form. It cannot discharge an obligation to obtain and record a customer's own declaration.
Autonomous agents
This is the argument that will be made most often against this paper over the next eighteen months, so it deserves a direct answer.
Agents will substantially reduce the handling cost of a request cycle. They can detect a gap the moment a submission lands rather than when an analyst reaches the queue, compose a specific and well-targeted request, carry full case context between rounds - eliminating the context-loss problem described in Section 04 - chase at the right interval, validate on arrival, and reconcile answers against everything already known. Applied well, that is a large reduction in the institutional labour pool and a meaningful reduction in elapsed time. We expect it, and it is worth doing.
What an agent cannot do is retrieve a fact that exists in no system. When the required input is the percentage of an LLC held by a member who is a natural person, and no register anywhere records it, the agent's only available action is to ask the owner. It will ask faster, more precisely and more politely than a human, and the owner will still have to stop what they are doing, find the answer, and reply.
What agents actually address
| Cost pool | Annual | Agent impact |
|---|---|---|
| Institutional labour | $1.35bn | Largely addressable |
| Merchant labour | $1.66bn | Partial: ask better, ask once |
| Deferred activation | $2.75bn | Marginal: merchant latency unchanged |
| Attributable abandonment | $4.65bn | Marginal |
Agents address roughly 13% of the modelled cost directly and can take a large share of it. Institutions should size the business case against that share rather than against the whole.
The economics of this are worth being precise about. Agents attack roughly 13% of the modelled cost, the institutional labour pool, and can plausibly take a large share of it. They attack merchant labour partially, by asking better and asking once. They do not directly reduce the merchant's latency, which is the dominant driver of elapsed time and therefore of abandonment. A world with excellent agents is a world where the same information is requested, from the same owner, with the same underlying delay, at lower cost to the institution. That is a good outcome. It is not a solution to the ceiling, and describing it as one will lead institutions to under-invest in the remedies that actually work.
Every remedy currently attracting investment operates on the retrieval side or the handling side. The residual is untouched by all of them. Institutions should still pursue orchestration and agents on their own merits. Vendor capability varies materially below the ceiling and in how efficiently requests are handled, and those differences are worth paying for. What no vendor can do by itself is create a standardised authoritative source for a fact that was never collected or never released, so the expected impact should be sized against the share of cost that sits in handling rather than against the whole.
What actually moves the ceiling
Two interventions move the ceiling itself, collecting more at source and releasing what is already collected, and both require legislation; a third, portable attestation, leaves the ceiling where it is and reduces how often institutions have to climb it, which makes it the only one available to the private sector today.
Raise the collection ceiling: collect more at source
This is the policy lever, and it is the highest-leverage remedy by a wide margin. The Canadian comparison establishes that it works. The CBCA regime moved some ownership-and-control attributes into a publicly retrievable layer for federal corporations with a single statutory change, without making the full record of individuals with significant control public. Quebec extended the principle further, to a public register that reaches foreign entities required to register there.
The obvious counter is cost to business, and it is a fair one. But the cost is already being paid. A US small business owner currently declares beneficial ownership repeatedly, to every institution it deals with, in a different format each time, with no record that persists and no reuse. Filing once, in a structured form, at the point of formation, is cheaper for the business than declaring it four times a year on demand. The debate has largely been framed as transparency against burden. It should be framed as filing once against attesting continuously.
The August 2026 final rule makes this the least available remedy in the United States for the foreseeable future, and the realistic pressure now comes from outside. The rule landed during the United States' mutual evaluation by the Financial Action Task Force, which treats beneficial ownership transparency as a priority standard, and the OECD's Global Forum has separately called on the United States to report on the availability of accurate and current beneficial ownership information. Institutions planning a three-year data strategy should treat international standards pressure, not domestic legislation, as the variable most likely to move the US collection ceiling, and should not plan on it moving soon.
Raise the distribution ceiling: release what is already collected
Several fields are already held by a state and simply not shared. This is the cheapest available remedy because the collection cost is sunk. Three concrete asks, in ascending order of political difficulty:
- Machine-readable access to state filings as standard. Uniform schema, documented API, defined update cadence. Roughly nine fields, already collected fifty-one times over, would become reliably retrievable rather than variably scraped.
- Workable regulated-entity access to beneficial ownership registers. A route for supervised institutions to reach beneficial ownership data for customer due diligence is a narrower and more defensible proposition than full public disclosure. Where such access exists in principle it is consent-based, conditional and narrow in scope, which limits its operational value; where a register is closed outright, as with the New York disclosures, no route exists at all. Making conditional access usable at onboarding volumes would resolve much of the distribution ceiling without reopening the transparency debate.
- An operating address field distinct from the registered agent address. Trivial to collect, currently absent almost everywhere, and the source of a substantial share of address-mismatch request cycles.
Make the residual reusable: portable attestation
The third remedy does not move the ceiling. It changes how many times each institution has to climb it, and it is the only one that can be built without legislation.
The waste in the current system is not that a business owner has to declare things about their business. Some of that is irreducible and appropriate. The waste is that the identical declaration is made from scratch to every institution, with no persistence, no reuse and no accumulated verification history. A merchant with four financial relationships has assembled the same document set four times, and none of the four institutions knows the other three exist.
The same declaration, four times
| Relationship | Fields collected | Format | Reused |
|---|---|---|---|
| Business bank account | 31 | PDF upload | No |
| Card acceptance | 31 | PDF upload | No |
| Working capital lender | 31 | PDF upload | No |
| Marketplace | 31 | PDF upload | No |
| Total | 124 | 31 underlying facts | Zero reuse |
A portable attestation layer would treat a merchant's declaration as a durable, owned artefact rather than a transient form submission: structured, timestamped, held under the business's control, accompanied by the evidence supplied, and presentable to the next institution with a verifiable history of when it was made and what corroboration it carried. The receiving institution still makes its own decision and still owns its own risk determination. What it does not do is start from zero.
A portable-attestation model is a combined legal, commercial, governance and technical problem, and it should not be presented as a build. Three conditions have to hold for it to be more than an idea:
- Reuse rights. Whether an institution may re-present data it collected, and whether a business may take its own attestation elsewhere, is governed by contract in the first instance, and then by privacy and data protection law, confidentiality obligations, record-retention rules, intellectual property rights in screening and scoring outputs, and each receiving institution's own independent compliance duties. Contract is the term most often negotiated and least often understood: data reuse language in master service agreements is the single highest-leverage commercial clause in this market, and it is routinely settled by people who have never considered what it forecloses.
- Structure at the point of first collection. A declaration captured as a PDF upload is not reusable by anything. The same declaration captured as structured fields with an evidence trail is reusable by everything. This is a design decision made once, usually without recognising its consequence.
- Provenance the receiving institution can rely on. The receiver needs to know who attested, when, under what verification, and what has changed since - otherwise reuse simply transfers risk without transferring confidence.
None of this raises the ceiling. All of it reduces the number of times the same thirty-one fields have to be extracted from the same owner by different parties asking the same questions in different formats. On the model in Section 05, halving the number of institutions that must collect from scratch would remove more cost than fully automating the institutional labour pool.
The interim discipline, available immediately
Institutions that cannot wait for legislation or infrastructure have four moves available this quarter, in order of return on effort:
- Carry case context between requests. The single largest recoverable inefficiency identified in this research, and a workflow fix rather than a data fix.
- Ask once, completely. Front-load every plausible attestation-layer requirement for the merchant's segment and risk band into the first request, accepting a longer initial form in exchange for fewer round trips. Elapsed time is dominated by round trips, not by form length.
- Separate the ceiling in the requirement set. Tier every required field explicitly. Fields below the ceiling should normally be prefilled or independently verified with a clear correction path, rather than collected from scratch; asking a merchant to type in a formation date you already hold is a self-inflicted request cycle.
- Instrument the four metrics in Section 07. Nothing improves before it is counted.
Policy raises the ceiling; portability reduces how often it has to be climbed; workflow discipline reduces the cost of each climb. Only the third is fully within any single institution's control, and it is where every institution should start - while recognising that it addresses the smallest of the three.
What to do, by role
Compliance and operations leaders should measure requests per onboarding before buying anything else; policymakers should recognise that a collection requirement removed does not remove the obligation, it relocates the cost onto small businesses and the institutions that serve them.
For compliance and operations leaders at PSPs, acquirers and banks
- Tier your requirement set before your next vendor conversation. Classify every required field by whether an independent source exists. The proportion above the ceiling is the hard bound on any automation business case, and knowing it changes what you should be buying.
- Report requests per onboarding to your executive committee monthly. It correlates with time to activation more strongly than any metric currently on that report.
- Stop measuring progress in straight-through-processing rate alone. It systematically excludes the population that determines your cycle time. Pair it with mean and 90th-percentile elapsed time including all request cycles.
- Split your funnel by eventual risk outcome. If applicants later flagged for risk complete your process at a higher rate than those never flagged, your friction is selecting against you, and you can establish that from data you already hold.
- Treat data reuse rights as a first-order commercial term. In partner, white-label and processor agreements, the clause governing whether collected business data can be re-presented is worth more than most of the pricing schedule.
- Fix context loss before buying anything. It is the cheapest material improvement available and it requires no new data.
For policymakers and supervisors
- An obligation without a source relocates cost; it does not remove it. Requiring institutions to identify beneficial owners while exempting entities from filing beneficial ownership does not reduce burden. It moves the burden from a one-time structured filing onto a repeated unstructured request, borne by small businesses, several times a year, for the life of the business.
- Distinguish collection from distribution in the design of any register. A register that is collected and closed imposes the full compliance cost on business while delivering none of the operational benefit. If access is politically constrained, constrain it to supervised institutions performing customer due diligence rather than closing it entirely.
- Machine readability is a substantive policy choice, not an implementation detail. A field published only as a scanned image or a per-record paid lookup is, for operational purposes, above the distribution ceiling.
- Consider measuring documentary attrition as a financial inclusion indicator. Applications abandoned during information exchange, split by firm age and size, would reveal an exclusion channel that no current supervisory metric captures.
- Note the asymmetry the Canadian comparison exposes. Two adjacent jurisdictions, one common merchant population, one shared obligation, and radically different data availability. Cross-border institutions are absorbing that difference as operating cost today.
The most valuable action available to almost every reader of this paper costs nothing and requires no vendor: count how many times you had to ask.
Method, limitations and definitions
What this research is
This paper is a structural analysis supported by a transparent economic model. It is not a survey. No practitioner panel was fielded, and no finding here should be cited as survey evidence. Where we state a number derived from our model, the inputs are published in Section 05 and can be replaced by any reader who disagrees with them.
Sources
- Regulatory position. FinCEN beneficial ownership reporting materials the interim final rule of 26 March 2025 and the final rule of 11 August 2026; New York Department of State guidance on the LLC Transparency Act and the December 2025 veto of the decoupling amendment; Corporations Canada guidance on individuals with significant control; Quebec's Legal Publicity Act regime. Positions stated as at August 2026; this is an actively moving area and readers should verify current status.
- Market structure. US Census Bureau Business Formation Statistics; SBA Office of Advocacy Small Business Profiles 2025; published financial crime compliance cost studies for the US and Canada.
- Operator data. Internal onboarding data shared with Detected under confidentiality by a global payment service provider, presented in abstracted form with absolute volumes and internal system identifiers withheld.
- Field map and tier assignment. Expert structural assessment by Detected, informed by production KYB and merchant onboarding implementations across acquiring, banking, lending and marketplace use cases.
Limitations we want stated plainly
- The requests-per-onboarding parameter is the weakest input in the model and the one that drives most of the variance. It is drawn from practitioner observation rather than measurement, because no measurement exists. Establishing it empirically is the single highest-value piece of follow-up research in this area.
- The field map reflects a composite requirement set. Individual institutions will differ, sometimes materially, particularly across risk bands and verticals.
- Tier assignment is made at the level of what governs the majority of onboarding events, not jurisdiction by jurisdiction. The full field list and its tier assignments are published in Table A.1 below so that any individual assignment can be challenged. A field-by-field, jurisdiction-by-jurisdiction matrix across all sixty-four filing jurisdictions would sharpen the estimate materially, and we have not built one.
- The single-operator case is illustrative of a pattern we observe widely. It is one institution, in one period, and is not presented as representative of the market.
- Abandonment attribution is inherently contestable. Applications end for many reasons, and we cannot cleanly isolate information requests as cause. The model's range is wide for this reason.
Definitions used throughout
| Term | Definition |
|---|---|
| Collection ceiling | The upper bound on available business data, set by what governments and institutions chose to collect. Introduced in this paper. |
| Distribution ceiling | The lower bound sitting beneath it, set by what is released, to whom, in what form, at what cadence. Introduced in this paper. |
| Attestation layer | The set of required data points with no generally available, independently retrievable authoritative source at commercial scale, obtainable only with the applicant's cooperation. Introduced in this paper. |
| Customer declaration | A statement the applicant makes about itself: ownership percentages, expected volume, business model, source of funds. |
| Customer-supplied evidence | An artefact the applicant produces: bank statements, processor statements, financial statements, supplier invoices, organisational charts. |
| Request loop | The asynchronous cycle of gap detection, request, merchant latency, response and re-review that governs onboarding elapsed time. Introduced in this paper. |
| Onboarding event | One SMB taken on by one regulated provider, requiring a KYB and due diligence assessment. |
| Tier 1 - 4 | Source classification for a required field: registry-derivable, other authoritative record, commercially inferable, attestation-only. |
| Perishability | The share of delayed trading activity permanently lost rather than deferred to a later period. |
Table A.1 - The full field map
| # | Field | Group | Tier | Source or source category | Access mode | Jurisdictional exceptions |
|---|---|---|---|---|---|---|
| 1 | Registered legal name | A | T1 | State or provincial register | Retrievable | – |
| 2 | Entity type | A | T1 | State or provincial register | Retrievable | – |
| 3 | Jurisdiction of formation | A | T1 | State or provincial register | Retrievable | – |
| 4 | Formation date | A | T1 | State or provincial register | Retrievable | – |
| 5 | Registry filing number | A | T1 | State or provincial register | Retrievable | – |
| 6 | Standing or good standing | A | T1 | State or provincial register | Retrievable | Update cadence varies widely |
| 7 | Registered agent | A | T1 | State or provincial register | Retrievable | Not applicable in all provinces |
| 8 | Registered office address | A | T1 | State or provincial register | Retrievable | Frequently not the operating address |
| 9 | Trading names and assumed-name filings | A | T2 | County, state or provincial assumed-name filings | Retrievable | Highly fragmented; often county level in the US |
| 10 | Federal tax identifier (EIN, BN) | B | T2 | IRS or CRA | Verify only | Value must be supplied before it can be checked |
| 11 | Tax identifier to legal-name match | B | T2 | IRS TIN matching or equivalent | Verify only | Access restricted by programme eligibility |
| 12 | Securities registration status | B | T2 | SEC EDGAR, SEDAR+ | Retrievable | Public issuers only; almost never applicable to SMBs |
| 13 | Sector licences and permits held | B | T2 | Federal, state and provincial licensing bodies | Retrievable | Sector-specific, no unified index |
| 14 | Directors and officers of record | C | T1 | State, provincial or federal register | Retrievable | Some US states only, often stale. Retrievable federally in Canada |
| 15 | Beneficial owners, identity | C | T4 | Applicant | Customer declaration | Partly retrievable for CBCA corporations and in Quebec |
| 16 | Ownership percentage per owner | C | T4 | Applicant | Customer declaration | Not captured by any US register |
| 17 | Ownership chain and intermediate entities | C | T4 | Applicant | Customer declaration | – |
| 18 | Control persons without equity | C | T4 | Applicant | Customer declaration | Description of control partly public under CBCA |
| 19 | Signatory authority for this application | C | T4 | Applicant | Customer declaration | – |
| 20 | Affiliated and related entities | C | T4 | Applicant | Customer declaration | Partially inferable from common officers where officers are filed |
| 21 | Nominee or trust arrangements | C | T4 | Applicant | Customer declaration | – |
| 22 | Owner identity documents | C | T4 | Applicant | Customer-supplied evidence | – |
| 23 | Website and owned digital estate | D | T3 | Open web, domain records | Inferable | Confidence low for newly formed firms |
| 24 | Operating or trading address | D | T4 | Applicant | Customer declaration | No US register holds an operating address distinct from registered office |
| 25 | Additional operating locations | D | T4 | Applicant | Customer declaration | – |
| 26 | Products and services actually sold | D | T4 | Applicant | Customer declaration | Fragments appear in licensing and sales-tax registrations |
| 27 | Business model and how revenue is earned | D | T4 | Applicant | Customer declaration | – |
| 28 | Fulfilment model and delivery timelines | D | T4 | Applicant | Customer declaration | – |
| 29 | Refund and cancellation policy | D | T4 | Applicant | Customer-supplied evidence | Sometimes published on the applicant's own site |
| 30 | Customer acquisition channels | D | T4 | Applicant | Customer declaration | – |
| 31 | Seasonality profile | D | T4 | Applicant | Customer declaration | Partly inferable from prior processing history where supplied |
| 32 | Bank account ownership confirmation | E | T2 | Account verification networks, open banking | Verify only | Requires an account value and customer consent |
| 33 | Expected annual processing volume | E | T4 | Applicant | Customer declaration | – |
| 34 | Expected average transaction value | E | T4 | Applicant | Customer declaration | – |
| 35 | Highest anticipated single transaction | E | T4 | Applicant | Customer declaration | – |
| 36 | Current or prior processor statements | E | T4 | Applicant | Customer-supplied evidence | – |
| 37 | Historic chargeback ratio | E | T4 | Applicant | Customer-supplied evidence | Held by card network programmes, not released to prospective acquirers |
| 38 | Bank account details | E | T4 | Applicant | Customer declaration | – |
| 39 | Bank statements | E | T4 | Applicant | Customer-supplied evidence | Obtainable via open banking with consent, which is still cooperation |
| 40 | Annual financial statements | E | T4 | Applicant | Customer-supplied evidence | Retrievable for public issuers and some non-North-American regimes |
| 41 | Interim year-to-date financials | E | T4 | Applicant | Customer-supplied evidence | – |
| 42 | Trade payment behaviour | F | T3 | Commercial bureaux and trade panels | Inferable | Thin or absent for young and small firms |
| 43 | Estimated firmographics (revenue, headcount) | F | T3 | Commercial data providers | Inferable | Modelled estimates; confidence low at the small end |
| 44 | Key suppliers and vendors | F | T4 | Applicant | Customer declaration | Partly observable in import records for goods importers |
| 45 | Inventory sourcing evidence | F | T4 | Applicant | Customer-supplied evidence | – |
| 46 | Third-party fulfilment partners | F | T4 | Applicant | Customer declaration | – |
| 47 | Sanctions screening, entity | G | T2 | Sanctions and watchlist providers | Verify only | Operates on a name already supplied |
| 48 | Litigation and judgments | G | T2 | Court records | Retrievable | Coverage and machine access vary by court |
| 49 | UCC filings and secured lending | G | T2 | State UCC filing offices, provincial PPSA registries | Retrievable | – |
| 50 | Bankruptcy and insolvency history | G | T2 | Federal courts, insolvency registers | Retrievable | – |
| 51 | Adverse media, entity | G | T3 | Media and adverse-media providers | Inferable | Sparse coverage of small private firms |
| 52 | Sanctions and PEP screening, owners | G | T4 | Applicant, then screening providers | Customer declaration | Screening is mature; input names are not retrievable in the US |
| 53 | Declared industry classification or MCC | G | T4 | Applicant | Customer declaration | Registry SIC and NAICS codes are self-selected and often stale |
| 54 | Source of funds and source of wealth | G | T4 | Applicant | Customer declaration | Required for higher-risk bands only |
Tier is assigned against a US-formed private company. Access mode distinguishes fields that are retrievable from an independent source, verify only where a source will confirm a value already supplied but will not return it, inferable from commercial data with a confidence score, and customer declaration or evidence where the applicant is the only practical source. Assignments are expert judgement and are published here so that any individual row can be challenged.
Follow-up research
Three studies would materially advance this field. None of them exists today, and we would welcome collaboration on any of them: a practitioner survey establishing requests per onboarding across institution type and merchant segment, which would replace the weakest input in our model with measurement; a full field-level coverage matrix across all sixty-four North American filing jurisdictions, which would turn the structural assessment in Section 03 into a per-jurisdiction reference; and a completion-by-risk-outcome study testing the adverse selection hypothesis in Section 06 against real portfolio data. We would rather state plainly that this work is outstanding than imply it sits behind the paper.
This paper may be cited and quoted with attribution to Detected. The model may be reproduced and re-run with different inputs; we ask only that modified figures are not attributed to us. Where a figure originates from our model rather than a published source, it is described as such in the text.