Research story · Computing, networks & measurement

Short paths,
uneven connections.

The number of nodes tells us how large a network sample is. Its connections reveal another part of the system: the routes available between nodes, the concentration of links and the evidence needed to explain how that structure arose.

Imagine two networks containing the same number of machines. In one, each machine connects only to its immediate neighbours around a ring. In the other, additional links join distant parts of that ring. The inventory is unchanged, but a message may have a much shorter route. Counting machines leaves an important engineering question open: how are they connected?

My 2018 paper with Marco Alberto Javarone, From Bitcoin to Bitcoin Cash: a network analysis, approached that question through three historical samples. Bitcoin was observed in April 2016; Bitcoin Cash in August and December 2017. The reported datasets contained 7,025, 963 and 1,454 nodes respectively. They were partial observations collected through connected machines and address discovery.

Two measurements of proximity

In an unweighted graph, shortest-path distance counts the fewest links needed to travel between two nodes. Clustering asks how densely the neighbours of a node are connected to one another. The quantities capture different aspects of proximity: local connectedness and the routes across the wider structure.

The small-world model of Duncan Watts and Steven Strogatz shows how strong local clustering can coexist with short characteristic paths. In their construction, rewiring some connections supplies shortcuts while much of the local structure survives. This is why the arrangement of a given set of links can matter as much as its size.

Our samples had reported average shortest paths of roughly 1.7–1.8 links and clustering coefficients of roughly 0.49–0.58. The paper treated these as suggestive of small-world behaviour and left comparison with matched random graphs for further investigation. These measurements describe the sampled graphs at their historical observation dates.

How a connection pattern might emerge

Network models provide candidate mechanisms. In Barabási and Albert’s preferential-attachment model, a growing network favours links to nodes that already have many connections. Success in attracting links can reinforce itself as new nodes arrive.

Bianconi and Barabási’s fitness-based framework also allows nodes to differ in their ability to compete for connections. A sufficiently attractive newcomer can gain ground despite arriving later. Fitness is a modelling parameter; applying it to machines requires an account of which resources or behaviours it represents.

Our paper proposed resource-dependent fitness as a possible explanation for uneven outgoing connectivity. It did not estimate node fitness or identify a causal effect of resources on connections.

What would strengthen the explanation?

An observed distribution and a model capable of producing it leave a statistical question: how well does that model fit, and which alternatives remain plausible? Clauset, Shalizi and Newman’s work on empirical power laws distinguishes parameter estimation, goodness-of-fit testing and comparison with competing distributions. A fitted curve alone cannot complete those steps.

That discipline applies here. The 2018 paper explicitly called for further goodness-of-fit analysis. Its proposed growth mechanism remains a hypothesis.

A follow-on study could separate three tasks. First, document what the observation method can see and what a recorded connection means. Second, collect comparable observations over time. Third, test whether measured changes in resources help explain changes in connectivity after accounting for alternative mechanisms. This is a proposed research design, rather than a report of new results.

Connect the graph to the question

The interpretation also needs a clear unit of analysis. A node-level graph cannot, by itself, identify the people or organisations operating its nodes. An assessment of independent control needs additional evidence about operators and the decisions being studied.

Likewise, a route containing few links leaves transmission and processing time unspecified. Turning graph distance into a performance assessment requires measurements of those delays, the workload and the relevant operating conditions. A security assessment must also say which failures or adversarial actions it considers.

The methodological lesson is to make each connection between evidence and claim explicit. Record the observed network, examine its structure, test explanations for that structure, and then ask which operational questions the measurements can support. A topology study is most useful when its boundaries remain visible.