NODE · LON-01|LONDON --:--:--
DAC | Digital Asset Claims

Knowledge centre · 6 min read

Reading a wallet clustering result

What a cluster is, which heuristics produce it, how it fails, and what it can and cannot say about control.

Distinct groups of related addresses with thin bridges between them

What a cluster actually is

A cluster is a set of addresses that observed behaviour suggests are controlled by one party. It is a hypothesis produced by rules, not a fact published by the network. No chain records ownership; ownership is inferred from how addresses are used.

This matters because clusters are frequently read as if they were account statements. They are not. A cluster is an argument, and like any argument it has premises that can be examined, and failure modes that can be tested.

The heuristics that build clusters

Common-input ownership is the oldest and strongest: on UTXO chains, when several inputs are spent together in one transaction, the spender normally controlled all of them. It is reliable in ordinary use and defeated deliberately by collaborative transactions.

Change identification groups an address with the transaction that funded it, by recognising which output returned the remainder of a spend. It depends on wallet conventions such as script type, output ordering and round-number amounts, and it degrades when a wallet is configured to avoid those tells.

Behavioural signals carry weight on account-based chains where common-input reasoning does not apply: consistent fee funding from one source, repeated interaction with the same contracts, address reuse, and timing patterns that hold across weeks rather than a single day.

Finally, service classification separates addresses that belong to exchanges, bridges and other infrastructure from addresses that behave like individual controllers. Missing that distinction is how a cluster ends up appearing to contain millions of unrelated users.

How clustering fails

Over-clustering merges parties that are not the same. It usually comes from misreading change, from treating an aggregation address as a personal one, or from chaining several weak inferences until a large group forms with no strong link anywhere inside it.

Under-clustering separates activity that belongs together, typically where careful coin control, separate wallets or privacy tooling suppress the usual signals. Under-clustering is the safer failure, and it is the one an evidence-led method should prefer when the signals are ambiguous.

Both failures are compounded by time. Heuristics that fit a 2016 wallet do not fit a 2026 one, and clusters built from mixed-era behaviour need their assumptions restated for each period rather than applied uniformly.

How to read a cluster you have been given

Ask four questions of any cluster. Which heuristics produced it? What is the strongest single link inside it, and what is the weakest? Does it contain any address classified as a service? And what confidence grade is attached to the cluster as a whole, as distinct from the individual transactions inside it?

A cluster presented without those answers should be treated as a lead, not as a finding. In a properly documented file, each of the four is already written down, and the grade sits on the cluster rather than being borrowed from the certainty of the underlying transaction data.

Be particularly careful with a cluster's edges. The core of a cluster is often well supported while its outermost members rest on one weak inference, and it is usually the outermost members that a reader most wants to be true.

What a cluster can and cannot support

A well-evidenced cluster can support statements about co-ordinated control of a set of addresses, about scale of activity, and about the relationship between one group of addresses and another. It can also usefully narrow further work by showing where corroboration is most likely to be found.

It cannot, on its own, name a person. Naming requires independent evidence outside the chain — registry records, published disclosures, service confirmations obtained lawfully — and the link between the cluster and that evidence must itself be documented and graded. A cluster plus a plausible name is not an attribution; it is a hypothesis with a name attached to it.

Continue reading

All guides