Proximity in the product space is the minimum of two conditional probabilities, so it cannot take a negative value and the repulsion half of the co-occurrence matrix has never been drawn. Drawing it with a margin-preserving null returns one hard fact and one hard limit. In 2024, of 190 exporting economies, the number holding a revealed advantage in both crude petroleum and women’s blouses is 0, against a chance expectation of 13.16. Conditioning that deficit on income group and port latitude removes its statistical significance, and this page stops where the conditioning stops it.
Instrument
Curveball fixed-fixed co-occurrence null (Strona et al. 2014, Nat Commun 5:4114), Benjamini-Hochberg tail at q = 0.01
Borrowed from
Community ecology: Diamond’s assembly rules, the Connor and Simberloff critique, the Stone and Roberts C-score
The product space is built from one half of a co-occurrence matrix. Hidalgo-Hausmann proximity is min(P(a|b), P(b|a)), a quantity bounded below by zero, so two headings that co-occur far less often than chance would predict are indistinguishable from two headings that merely never meet. This page computes the other half: for every pair of HS4 headings held by enough economies to test, the observed number of joint holders against the distribution generated by a curveball fixed-fixed null, which preserves every economy’s diversity and every heading’s ubiquity and randomises nothing else.
One naming rule governs everything below, and it is not cosmetic. The measured object is statistical incompatibility, defined as a co-occurrence deficit against a margin-preserving null. It is not competition and it is not exclusion in any causal sense. Where the parquet columns say exclusion and attraction, read deficit and surplus. Community ecology spent two decades on exactly this ambiguity and did not resolve it; nothing in a trade matrix resolves it either, and the habitat control below shows how much of the deficit the two obvious confounds absorb.
The estimator is checked before it is trusted. The known attraction structure of the product space, the apparel cluster, has to come back out of the null first (Figure 1). It does. Only then is anything negative reported.
The control first: the null recovers the attraction structure it was not told about
The negative
The two panels below are the same matrix, the same null and the same significance rule, read in opposite directions. Nodes are HS4 headings placed on the circle in code order, so each HS section owns a contiguous arc and a chord that crosses the disc is a link between distant parts of the classification. The right panel is the product space the site already ships: a dense apparel knot, links short, everything inside one section. The left panel is its negative, and it is not a mirror image. It is a small number of hubs, crude petroleum above all, wired to the far side of the classification.
Figure 2
The anti-product space and the product space, 2024
Deficit: 250 pairs, 184 headings
Figure 3
The same negative at four points in the panel
1995: 1,317 significant, 150 drawn
Crude petroleum and women’s blouses
The pair the proposal was built on is HS4 2709, crude petroleum, against HS4 6206, women’s and girls’ blouses and shirts, not knitted. In 2024 the intersection is empty: 37 economies hold the first, 42 hold the second, 0 hold both, against a null expectation of 13.16 (sd 2.5724), z = -5.12, p = 3.12 x 10^-7.
The proposal also said no economy had held both for thirty years. That is false, and it is the largest correction on this page. In 1995 eight economies held both (Netherlands Antilles, Egypt, Indonesia, Latvia, Oman, Syria, Tunisia, Vietnam), and at least one economy held both in every year from 1995 to 2023. Albania held both in every year from 2010 to 2022, alone in most of them. Curaçao is the last, in 2023. The defensible claim is the decline and its significance, not an absence: the count falls from 8 to 0 while the null expectation moves only from 18.18 to 13.16, and z is negative in all 30 years, from -3.64 to -5.40, BH-significant in every one of them.
A second pair does the same thing without reaching zero until the same year: crude petroleum against HS4 6107, men’s and boys’ knitted underwear and nightwear. In 1995 it was observed 8 times against 16.30 expected, z = -2.99, which does not clear BH at q = 0.01. In 2024 it is observed 0 times against 10.69, z = -4.52, p = 6.17 x 10^-6, and it does. A pair can enter this tail without ever having been in it.
Figure 4
2709 × 6206: observed joint holders against the curveball null, 1995 to 2024
Year
Observed
Figure 5
The tail, year by year
Figure 6
The twelve strongest co-occurrence deficits, 2024
Pair
Heading a
Heading b
Ub. a
Ub. b
Obs.
Exp.
z
2709 × 6211
Petroleum oils and oils obtained from bituminous mine…
Track suits, swimwear and other garments (not knitted…
37
50
0
15.11
-6.54
2709 × 4819
Petroleum oils and oils obtained from bituminous mine…
Cartons, boxes, cases, bags and the like, of paper, p…
37
54
2
16.14
-5.65
2709 × 6203
Petroleum oils and oils obtained from bituminous mine…
Packing cases, boxes, crates, drums and similar packi…
Gold (including gold plated with platinum) unwrought …
45
49
2
17.80
-5.59
2202 × 2709
Waters, including mineral and aerated waters, contain…
Petroleum oils and oils obtained from bituminous mine…
65
37
4
18.93
-5.59
2710 × 6115
Petroleum oils and oils from bituminous minerals, not…
Hosiery; panty hose, tights, stockings, socks and oth…
55
28
1
12.92
-5.49
5202 × 7204
Cotton waste (including yarn waste and garnetted stoc…
Ferrous waste and scrap; remelting scrap ingots of ir…
The habitat control, and what it costs
A fixed-fixed null preserves margins and nothing else. If oil economies and low-wage garment economies are disjoint country sets for endowment reasons, that shows up as a co-occurrence deficit whether or not the two capabilities are incompatible. The two obvious confounds are income and climate, so both are conditioned on directly: the whole estimator is re-run inside each income group and each port-latitude band, and separately the tail is described by how far apart the two headings’ holder sets are on each covariate.
Both tests cut against the strong reading, and they are reported that way. Inside strata the deficit tail almost vanishes: 9 significant pairs out of 320,400 tested inside High income, 0 out of 116,886 inside Upper middle income, and none at all outside the tropics on the latitude split. The headline pair keeps its direction in every stratum, at 0 observed against expectations of 0.54 to 7.09, but z falls to between -4.06 and -0.89 and it is BH-significant in none of them. Part of that is power loss and part of it is real habitat filtering. This data cannot separate the two.
Figure 7
Re-running the whole estimator inside income groups and latitude bands
Stratum
Economies
Ub. floor
Pairs tested
Deficit tail
Surplus tail
2709 × 6206 obs.
exp.
z
Control z
All economies
190
25
93,096
1,418
2,618
0
13.19
-5.37
+7.95
High income
71
9
320,400
9
24
0
3.68
-2.77
+3.17
Upper middle income
45
6
116,886
0
0
0
4.05
-3.26
+3.68
Lower middle income
46
6
56,280
4
56
0
3.58
-3.24
+4.57
Low income
24
5
3,160
0
0
0
0.54
-0.89
+3.78
Tropical (|lat| < 23.5)
82
11
25,425
37
287
0
7.09
-4.06
+6.00
Subtropical (23.5 to 45)
45
6
292,230
0
1
0
4.27
Figure 8
The deficit tail concentrates exactly where the confounds predict
Decile
Income gap
% in deficit tail
mean z
Latitude gap (deg)
% in deficit tail
mean z
1
0.035
0.99%
+0.499
0.57
0.86%
Basket tension
The pair statistic aggregates to economies. Basket tension is the mean z over every tested pair inside an economy’s own RCA >= 1 set: negative means the things it exports are things the null says rarely travel together. It is a description of a basket, not a diagnosis of it, and it inherits every confound above.
Figure 9
Basket tension, 2024
Figure 10
Entry advisory: the most incompatible target for each economy
Every free choice, swept
Nine settings can be moved: the number of null replicates, the null seed, the mixing length, the Balassa threshold, the ubiquity floor, the BH level, the exporter size floor, the HS-revision handling rule and the product grain. Each gets its own figure, because a number that only exists at one setting is not a result. Two of these sweeps change the headline, and both changes are printed rather than resolved in the estimator’s favour.
Figure 11
Replicate count: the tail size has not converged at the panel setting
Replicates
Deficit tail
Surplus tail
Headline expected
Headline sd
Headline z
25
2,536
3,240
12.560
1.873
-6.71
50
1,851
Figure 12
Null seed: how much of the tail is one draw
Seed
Deficit tail
Surplus tail
Share of tested
Headline expected
Headline sd
Headline z
11
1,438
2,555
1.54%
13.17
2.468
-5.34
777
1,482
2,606
1.59%
12.70
2.524
-5.03
12345
1,398
2,569
1.50%
13.03
2.410
-5.41
99991
1,409
2,557
1.51%
13.32
2.439
-5.46
20260802
1,397
2,622
1.50%
13.16
2.572
-5.12
Five independent null chains at 200 replicates give deficit tails of 1,397 to 1,482 pairs, a spread of 6.1%. The headline pair’s expectation moves between 12.70 and 13.32 and its z between -5.46 and -5.03.
Seed noise is on top of the replicate bias in Figure 11, not instead of it: a tail count from this page carries up to 6.1% of seed noise across five chains and, at 200 replicates, about 8.7% of upward bias. Every other figure uses seed 20260802. Sweep runs use seed 20260802 + 7000, which is why sweep tail counts differ from the panel by a few percent even at identical settings.
Seed sweep: data/parquet/instruments_anti_product_space_sweep_seed.parquet (2024, 200 replicates per chain).
Figure 13
Balassa threshold: the zero is a fact about RCA >= 1
RCA threshold
Economies
Links
Pairs tested
Deficit tail
2709 × 6206 obs.
exp.
z
Control obs.
Control z
0.50
190
42,447
371,953
11,308
3
21.84
-6.44
47
+8.38
0.75
190
33,831
193,131
4,109
0
15.13
-5.91
43
+8.23
1.00
190
28,049
93,096
1,475
0
13.18
-5.22
35
+8.67
1.25
190
23,959
47,586
335
0
10.02
-4.46
30
+8.39
1.50
190
20,803
27,966
99
0
8.95
-4.07
28
+8.02
2.00
190
16,401
7,140
16
0
6.67
-3.35
25
+9.27
3.00
190
11,395
1,081
1
0
3.87
-2.40
Figure 14
Ubiquity floor: which pairs are allowed into the test at all
Ubiquity floor
Pairs tested
Deficit tail
Surplus tail
Share of tested
BH p cutoff
Headline testable
5
760,761
921
7,381
0.12%
1.09 x 10^-4
yes
10
654,940
974
6,987
0.15%
1.21 x 10^-4
yes
15
431,056
1,075
5,515
0.25%
1.53 x 10^-4
yes
20
202,566
1,299
3,922
0.64%
2.58 x 10^-4
yes
25
93,096
1,397
2,622
1.50%
4.31 x 10^-4
yes
30
38,226
1,033
1,234
2.70%
5.92 x 10^-4
yes
40
6,786
265
303
3.91%
8.34 x 10^-4
no
50
703
55
27
7.82%
1.10 x 10^-3
no
75
6
0
3
0.00%
2.26 x 10^-4
no
100
0
0
0
Figure 15
False-discovery level: the tail is a choice of q
BH q
Deficit tail
Surplus tail
Share of tested
BH p cutoff
0.0001
50
505
0.05%
5.87 x 10^-7
0.001
226
1,035
0.24%
1.35 x 10^-5
0.005
819
1,938
0.88%
1.48 x 10^-4
0.01
1,397
2,622
1.50%
4.31 x 10^-4
0.05
5,306
5,267
5.70%
5.68 x 10^-3
0.1
9,151
7,436
9.83%
1.78 x 10^-2
At q = 0.0001 the deficit tail is 50 pairs; at q = 0.1 it is 9,151. The surplus tail is larger than the deficit tail at every level tested.
q = 0.01 is used everywhere else on this page. The headline pair survives every level tested, with p = 3.12 x 10^-7 against a BH cutoff of 4.31 x 10^-4 at the chosen level.
SELECT value, n_sig_exclusion, n_sig_attraction, share_sig_exclusion, bh_p_cutoff
FROM read_parquet('data/parquet/instruments_anti_product_space_sweep_q.parquet') ORDER BY value
Figure 16
Exporter size floor: which economies are in the matrix
Size floor
Economies
Links
Pairs tested
Deficit tail
2709 × 6206 obs.
exp.
z
$10M
219
29,847
118,341
1,425
0
12.90
-5.09
$50M
199
28,519
97,020
1,560
0
13.39
-5.56
$100M
190
28,049
93,096
1,475
0
13.18
-5.22
$500M
167
26,626
78,210
1,245
0
12.53
-5.40
$1B
157
26,194
73,920
1,021
0
12.52
-5.14
$5B
123
23,285
36,315
542
0
9.95
-5.17
Moving the floor from $10M to $5B moves the matrix from 219 economies to 123. The headline pair is observed 0 times at every floor tested, with z between -5.56 and -5.09.
The floor of $100M used throughout removes economies whose RCA vector is dominated by a handful of shipments. Values are millions of USD, converted from BACI thousands.
Mixing: how far the curveball walks between replicates
Trades per replicate
Swaps
Deficit tail
Surplus tail
Headline expected
Headline z
0.5E
14,024
1,472
2,593
13.06
-5.31
1E
28,049
1,393
2,533
12.99
-4.95
2E
56,098
1,475
2,599
13.18
-5.22
5E
140,245
1,463
2,584
13.29
-5.51
10E
280,490
1,411
2,585
12.89
-5.35
From half a swap per link to ten, the deficit tail moves between 1,393 and 1,475 pairs and the headline expectation between 12.89 and 13.29. Mixing is already adequate at 0.5E.
E is the number of links in the matrix (28,049 in 2024), so 2E trades per replicate is 56,098 attempted curveball swaps between draws. This is the setting that would show up as a stuck chain if it were too small, and it does not.
SELECT value, trades_per_replicate, n_sig_exclusion, n_sig_attraction,
exp_2709x6206, z_2709x6206 FROM read_parquet('data/parquet/instruments_anti_product_space_sweep_trades.parquet') ORDER BY value
Figure 18
HS-revision handling: the specification that double counts, and the two that do not
Spec
Rule
Economies
Headings
Links
World exports
Pairs tested
Deficit tail
Headline obs.
exp.
z
Control obs.
A
base + ext deduplicated to latest revision, ext-only headings
190
1,242
28,049
$23.273 trillion
93,096
1,475
0
13.18
-5.22
35
B
naive UNION ALL over every ext revision (double counts)
196
1,242
31,537
$48.350 trillion
147,153
2,232
0
14.99
-5.82
38
C
base HS92 catalog only
190
1,217
27,436
$22.832 trillion
89,253
1,292
0
12.99
-5.32
35
A plain UNION ALL over the revision partitions of country_year_product_ext counts the same trade up to six times and returns $48.350 trillion of world exports for 2024, against $23.273 trillion deduplicated and $22.832 trillion from the HS92 base alone. That inflation is what pushes 6 extra small economies over the size floor, and it is where a “196 economies” figure comes from.
country_year_product_ext stores the same HS6 code once per HS revision that contains it, so 010190 appears under HS02, HS07, HS12, HS17 and HS22 with different values. Summing across revisions is double counting, not a modelling choice, which is why the primary specification deduplicates to the latest revision per country-product. The headline pair is observed 0 times under all three specifications; its expectation ranges 12.99 to 14.99 and its z -5.22 to -5.82. Quote the range, not one end of it.
Figure 19
Product grain: the result does not survive chapter aggregation
Grain
Products
Links
Pairs tested
Deficit tail
Share of tested
Replicates
Headline co-occurrence
HS2 chapter
96
3,676
2,628
72
2.74%
200
chapter 27 × 62: 4
HS4 heading
1,242
28,049
93,096
1,475
1.58%
200
2709 × 6206: 0
HS6 subheading
6,434
116,429
990,528
18,488
1.87%
50
max 2709xx × 6206xx: 0
At HS2, mineral fuels (chapter 27) and non-knitted apparel (chapter 62) co-occur in 4 economies, so the deficit is destroyed by aggregation. At HS6 the largest co-occurrence between any 2709xx subheading and any 6206xx subheading is 0. The claim lives at HS4 and at HS6, and not above it.
The HS6 run uses 50 replicates rather than 200 (990,528 tested pairs), so its tail count is inflated relative to the HS4 figure by more than the bias in Figure 11. The share of tested pairs in the deficit tail is nevertheless close across grains: 2.74% at HS2, 1.58% at HS4, 1.87% at HS6.
SELECT value, grain_label, n_headings, n_links, n_pairs_tested, n_sig_exclusion,
share_sig_exclusion, replicates, obs_27x62, obs_2709x6206, obs_2709x6206_hs6_max
FROM read_parquet('data/parquet/instruments_anti_product_space_sweep_grain.parquet') ORDER BY value
What the proposal said, and what the computation returned
“196 exporting economies.” Cut. 196 economies and 31,537 links are the counts under the double-counted specification B. Under the deduplicated primary specification A it is 190 economies and 28,049 links; on the HS92 base alone (specification C), 190 and 27,436.
“Ubiquity 41 and 45, control 38.” Cut. Those are specification B values. Under the primary specification: 37 and 42, control 35.
“A chance expectation near fifteen.” Cut. The curveball expectation is 13.16 under the primary specification (12.70 to 13.32 across five null seeds) and 12.99 on the base catalog. 14.99 is specification B, which is where fifteen came from.
“And none has for thirty years.” False. Eight economies held both in 1995 and at least one held both in every year through 2023. The intersection is empty in 2024 only.
“About 1,400 pairs in the tail.” Corrected to about 1,300: 1,285 at 1,000 replicates against 1,397 at the panel setting of 200 (Figure 11), plus up to 6.1% of seed noise on top (Figure 12).
The aggregate C-score excess. Cut entirely. It was weak enough that a reader stopping at the summary statistic would correctly call it noise. What remains of the aggregate is descriptive only and is printed in the note to Figure 5.
What replicated exactly, checked twice end to end and once more by an independent SQL-only path with no numpy: the 2024 matrix (190 economies, 1,242 headings, 28,049 links), the ubiquities of all five headline headings, the empty intersections of 2709 × 6206 and 2709 × 6107, and the positive control at 35 joint holders. Every null-derived number is bit-identical between the two full runs; the only differences anywhere are the last two digits of world exports in two years, a float summation-order artefact of DuckDB’s parallel aggregate at relative size about 1e-16.
The strongest objection to this page
A fixed-fixed null preserves margins only, so structure induced by a third variable shows up as exclusion when it is really shared habitat filtering. Oil states and low-wage garment states are disjoint country sets for endowment reasons that no co-occurrence null can separate from capability incompatibility. This is the unresolved habitat-filtering versus competition problem ecology never settled either.
Conceded, and the claim was changed to fit it rather than defended. The measured object is named statistical incompatibility throughout and defined as a co-occurrence deficit against a margin-preserving null; the words competition and exclusion do not appear unqualified anywhere on this page.
The runs that answer it are Figures 7 and 8. Conditioning on the two obvious confounds does not rescue the strong reading, it damages it. Inside income groups and latitude bands the deficit tail nearly disappears: 9 significant pairs of 320,400 tested inside High income, 0 of 116,886 inside Upper middle income, none outside the tropics on the latitude split. The headline pair keeps its sign in every stratum, still 0 observed, but its z falls to between -4.06 and -0.89 and it is BH-significant in none of them. At the pair level, the deficit tail is 4.4 times as dense in the top income-gap decile as in the bottom, and 5.4 times as dense in the top latitude-gap decile.
The promotion rule was that only pairs surviving both conditionings reach the headline table. Applied honestly, that rule empties the table: the headline pair does not clear BH inside any stratum. Part of the collapse is genuine power loss (fewer economies, a noisier null standard deviation, and a floor scaled to stratum size that enlarges the tested set and hardens the BH correction), and part of it is real habitat filtering. Nothing in this data separates the two, exactly as the objection says, and the page does not pretend otherwise.
So what survives is narrower than the proposal: a descriptive fact under a stated convention (0 joint holders in 2024 at RCA >= 1, 3 at RCA >= 0.5), a direction that is present in every stratum, and a tail whose composition is consistent with habitat filtering. The strong reading, that these capabilities are incompatible, is not available from this page.
Every one of the ten strongest positive pairs in 2024 is an apparel pair, and the strongest, 6102 × 6110, is observed 36 times against 13.72 expected. The estimator reproduces the product space before it is asked to invert it.
z is a standardised co-occurrence surplus, not a t-statistic: (observed minus null mean) divided by the null standard deviation over 200 curveball replicates. The p-value is two-sided normal on that z, which is why the extreme tail values are reported as exponents rather than as zero. All ten pairs clear Benjamini-Hochberg at q = 0.01 (the BH p cutoff for 2024 is 4.31 x 10^-4).
CEPII BACI V202501, data/parquet/country_year_product/ and country_year_product_ext/ (export_value in THOUSANDS of USD; every musd column here is millions of USD, thousands / 1000). Primary specification A: HS92 base catalog plus country_year_product_ext deduplicated to the latest HS revision per country-product, ext restricted to HS4 headings absent from the base. Pair panel: data/parquet/instruments_anti_product_space_pairs.parquet (top 100 positive pairs per panel year).Query
SELECT hs_a, hs_b, name_a, name_b, ubiquity_a, ubiquity_b, observed,
null_mean, null_sd, z, p_two_sided FROM read_parquet('data/parquet/instruments_anti_product_space_pairs.parquet')
WHERE year = 2024 AND side = 'attraction' ORDER BY z DESC LIMIT 10
Textiles (44)
Vegetable Products (21)
Metals (19)
Foodstuffs (15)
Animal Products (11)
Machines (11)
Mineral Products (10)
Chemical Products (7)
Wood Products (7)
Paper Goods (7)
Plastics and Rubbers (5)
Miscellaneous (5)
Animal Hides (4)
Stone and Glass (4)
Transportation (4)
Precious Metals (3)
Instruments (3)
Footwear and Headwear (2)
Animal and Vegetable Fats and Oils (1)
Arts and Antiques (1)
Surplus: 100 pairs, 47 headings
Textiles (24)
Machines (7)
Chemical Products (3)
Metals (3)
Foodstuffs (2)
Mineral Products (2)
Animal Hides (2)
Miscellaneous (2)
Paper Goods (1)
Transportation (1)
The deficit panel is dominated by a few hubs: 250 of the 1,397 significant negative pairs are drawn here and they touch only 184 headings, with mineral fuels and precious metals on one side and textiles, wood and paper on the other. The surplus panel touches 47 headings and barely leaves one section.
Left panel: the 250 most negative BH-significant pairs of 1,397 in 2024. Right panel: the 100 most positive of 2,622. Both are truncated, in the precompute (to 400 and 100 per year) and again here, so the panels show the extremes, not the whole tail. Chord opacity and width scale with |z| within the panel; angular position is the HS4 code and carries no fitted information. Node area scales with the number of links the node has inside the panel.
CEPII BACI V202501, data/parquet/country_year_product/ and country_year_product_ext/ (export_value in THOUSANDS of USD; every musd column here is millions of USD, thousands / 1000). Primary specification A: HS92 base catalog plus country_year_product_ext deduplicated to the latest HS revision per country-product, ext restricted to HS4 headings absent from the base. Pair panel: data/parquet/instruments_anti_product_space_pairs.parquet (year, side, hs_a, hs_b, z).Query
-- left panel
SELECT hs_a, hs_b, z FROM read_parquet('data/parquet/instruments_anti_product_space_pairs.parquet') WHERE year = 2024 AND side = 'exclusion' ORDER BY z LIMIT 250;
-- right panel
SELECT hs_a, hs_b, z FROM read_parquet('data/parquet/instruments_anti_product_space_pairs.parquet') WHERE year = 2024 AND side = 'attraction' ORDER BY z DESC LIMIT 100;
2005: 702 significant, 150 drawn
2015: 1,112 significant, 150 drawn
2024: 1,397 significant, 150 drawn
The shape is stable across thirty years: a mineral-fuels hub wired to textiles, wood and paper. The tail it is drawn from is not stable in size, running from 1,317 significant negative pairs in 1995 to 1,397 in 2024 (Figure 5).
Each panel is the 150 most negative BH-significant pairs of that year, on the same layout rule. Panel years only: the pair-level extract was precomputed for 1995, 2005, 2015 and 2024. Labels are suppressed at this size; the colour key of Figure 2 applies.
CEPII BACI V202501, data/parquet/country_year_product/ and country_year_product_ext/ (export_value in THOUSANDS of USD; every musd column here is millions of USD, thousands / 1000). Primary specification A: HS92 base catalog plus country_year_product_ext deduplicated to the latest HS revision per country-product, ext restricted to HS4 headings absent from the base. Pair panel: data/parquet/instruments_anti_product_space_pairs.parquet.Query
SELECT year, hs_a, hs_b, z FROM read_parquet('data/parquet/instruments_anti_product_space_pairs.parquet') WHERE side = 'exclusion' AND year = ? ORDER BY z LIMIT 150;
Expected
sd
z
BH q = 0.01
Holders
1995
8
18.18
2.662
-3.83
significant
Netherlands Antilles, Egypt, Indonesia, Latvia, Oman, Syria, Tunisia, Vietnam
1996
7
15.97
2.462
-3.64
significant
Egypt, Indonesia, Latvia, Oman, Syria, Tunisia, Vietnam
1997
5
16.60
2.656
-4.37
significant
Indonesia, Latvia, Syria, Tunisia, Vietnam
1998
6
17.64
2.740
-4.25
significant
Egypt, Indonesia, Latvia, Syria, Tunisia, Vietnam
1999
6
17.76
2.722
-4.32
significant
United Arab Emirates, Indonesia, Latvia, Syria, Tunisia, Vietnam
2000
5
17.12
2.640
-4.59
significant
Brunei Darussalam, Indonesia, Syria, Tunisia, Vietnam
2001
4
16.05
2.687
-4.49
significant
Indonesia, Latvia, Syria, Vietnam
2002
4
16.93
2.769
-4.67
significant
Indonesia, Tunisia, Tuvalu, Vietnam
2003
3
16.60
2.536
-5.36
significant
Indonesia, Tunisia, Vietnam
2004
4
16.34
2.636
-4.68
significant
Brunei Darussalam, Indonesia, Tunisia, Vietnam
2005
3
15.52
2.643
-4.74
significant
Syria, Tunisia, Vietnam
2006
2
14.52
2.558
-4.89
significant
Syria, Vietnam
2007
3
14.41
2.487
-4.59
significant
Syria, Tunisia, Vietnam
2008
3
13.62
2.519
-4.22
significant
Syria, Tunisia, Vietnam
2009
4
13.54
2.284
-4.18
significant
Georgia, Syria, Tunisia, Vietnam
2010
3
15.89
2.426
-5.32
significant
Albania, Syria, Tunisia
2011
3
14.98
2.572
-4.66
significant
Albania, Syria, Tunisia
2012
2
14.06
2.385
-5.06
significant
Albania, Tunisia
2013
2
14.81
2.501
-5.12
significant
Albania, Tunisia
2014
1
14.89
2.639
-5.26
significant
Albania
2015
1
14.45
2.514
-5.35
significant
Albania
2016
1
14.08
2.421
-5.40
significant
Albania
2017
1
12.70
2.441
-4.79
significant
Albania
2018
2
12.09
2.409
-4.19
significant
Albania, Saint Lucia
2019
1
12.29
2.340
-4.83
significant
Albania
2020
1
14.13
2.601
-5.05
significant
Albania
2021
1
11.97
2.519
-4.36
significant
Albania
2022
1
11.63
2.501
-4.25
significant
Albania
2023
1
11.91
2.450
-4.45
significant
Curaçao
2024
0
13.16
2.572
-5.12
significant
none
The observed line runs below the null in every year of the panel and reaches zero once, in 2024. The null expectation is itself falling, from 18.18 to 11.63 at its lowest, because both headings are held by fewer economies than they used to be.
Ubiquity of 2709 falls from 39 economies in 1995 to 37 in 2024; ubiquity of 6206 from 53 to 42. The null holds both of those margins fixed, so the falling expectation is a consequence of the data, not of the test. The 2024 zero is a fact about the Balassa RCA >= 1 convention: at a threshold of 0.50 the pair has 3 joint holders (Figure 13).
CEPII BACI V202501, data/parquet/country_year_product/ and country_year_product_ext/ (export_value in THOUSANDS of USD; every musd column here is millions of USD, thousands / 1000). Primary specification A: HS92 base catalog plus country_year_product_ext deduplicated to the latest HS revision per country-product, ext restricted to HS4 headings absent from the base. Pair series: data/parquet/instruments_anti_product_space_headline_pairs.parquet; holder lists: data/parquet/instruments_anti_product_space_headline_holders.parquet.Query
SELECT year, count(*) AS n_dual, string_agg(iso3, ' ' ORDER BY iso3) AS holders FROM (
SELECT year, iso3 FROM read_parquet('data/parquet/instruments_anti_product_space_headline_holders.parquet') WHERE hs4 = '2709'
INTERSECT
SELECT year, iso3 FROM read_parquet('data/parquet/instruments_anti_product_space_headline_holders.parquet') WHERE hs4 = '6206'
) GROUP BY year ORDER BY year
The deficit tail is not a residue of the surplus tail. In 1995 it was the larger of the two (1,317 pairs against 888); by 2024 the surplus tail is roughly twice its size (2,622 against 1,397), on 93,096 pairs tested.
The tested set changes size across years because it is defined by a fixed ubiquity floor of 25 on a moving matrix, so counts are not directly comparable across years without the denominators, which are in the CSV. The aggregate direction of the whole matrix is deliberately not the headline: 58.52% of tested pairs have z < 0 in 2024 against 64.14% in 1995, and the mean z is -0.2320. A reader who stopped at that summary statistic would correctly call it noise. The claim is the tail, and only the tail.
CEPII BACI V202501, data/parquet/country_year_product/ and country_year_product_ext/ (export_value in THOUSANDS of USD; every musd column here is millions of USD, thousands / 1000). Primary specification A: HS92 base catalog plus country_year_product_ext deduplicated to the latest HS revision per country-product, ext restricted to HS4 headings absent from the base. Panel: data/parquet/instruments_anti_product_space_summary.parquet (30 years, specification A). Matrix: Balassa RCA >= 1, exporters above $100M total exports, HS4 grain. Null: curveball fixed-fixed swap (Strona et al. 2014, Nat Commun 5:4114), 200 replicates, 2E trades per replicate, seed 20260802. Tail: Benjamini-Hochberg at q = 0.01 over pairs where both ubiquities are at least 25.
Query
SELECT * FROM read_parquet('data/parquet/instruments_anti_product_space_summary.parquet') ORDER BY year
32
87
7
20.34
-5.49
6206 × 7108
Blouses, shirts and shirt-blouses; women's or girls' …
Gold (including gold plated with platinum) unwrought …
42
49
3
16.76
-5.48
2710 × 4421
Petroleum oils and oils from bituminous minerals, not…
Wooden articles n.e.c. in heading no. 4414 to 4420
55
28
2
12.41
-5.48
2709 × 6104
Petroleum oils and oils obtained from bituminous mine…
Petroleum oils and oils from bituminous minerals, not…
Coats; women's or girls' overcoats, carcoats, capes, …
55
32
1
14.36
-5.45
2709 × 6101
Petroleum oils and oils obtained from bituminous mine…
Coats; men's or boys' overcoats, car-coats, capes, cl…
37
44
0
13.76
-5.44
The strongest is 2709 × 6211, crude petroleum against track suits and swimwear: 0 observed against 15.11 expected, z = -6.54. Seven of the twelve have crude or refined petroleum on one side. The headline pair is not the extreme case, it is the legible one.
212 tested pairs have zero observed co-occurrence in 2024 and 153 of those clear BH at q = 0.01. A zero is not automatically significant: it depends on how ubiquitous both headings are. Nothing here is evidence of a causal mechanism, and the habitat control in Figure 8 is a condition of reading this table at all.
CEPII BACI V202501, data/parquet/country_year_product/ and country_year_product_ext/ (export_value in THOUSANDS of USD; every musd column here is millions of USD, thousands / 1000). Primary specification A: HS92 base catalog plus country_year_product_ext deduplicated to the latest HS revision per country-product, ext restricted to HS4 headings absent from the base. Pair panel: data/parquet/instruments_anti_product_space_pairs.parquet.Query
SELECT hs_a, hs_b, name_a, name_b, ubiquity_a, ubiquity_b, observed,
null_mean, null_sd, z, p_two_sided FROM read_parquet('data/parquet/instruments_anti_product_space_pairs.parquet')
WHERE year = 2024 AND side = 'exclusion' ORDER BY z LIMIT 12
-3.23
+3.94
High latitude (45+)
20
5
171,991
0
0
0
0.78
-1.12
+2.50
No maritime port
43
6
56,953
0
19
0
0.85
-1.12
+3.60
The full-sample run finds 1,418 significant deficits. Every stratified re-run finds between 0 and 37. The tail is largely a between-stratum phenomenon.
The all-economies row is a separate null chain from the panel run, which is why it reads 1,418 against the panel’s 1,397 at identical settings: that gap is seed noise, and its size is the subject of Figure 12. Specialisations are held fixed at their world definition; only the set of economies changes. The ubiquity floor is scaled to stratum size (25 of 190 economies, 13.2%, minimum 5), which makes the tested set inside a stratum much larger than the full-sample tested set and the BH correction correspondingly harsher: that is a mechanical part of the collapse and is not evidence for the null. The positive control survives every stratum (15 observed against 5.42 in the tropical band, z = +6.00), which is what says the estimator still has some power inside a stratum. Income groups are the World Bank classification carried in gmd.parquet; four economies are unclassified and appear only in the all-economies row.
CEPII BACI V202501, data/parquet/country_year_product/ and country_year_product_ext/ (export_value in THOUSANDS of USD; every musd column here is millions of USD, thousands / 1000). Primary specification A: HS92 base catalog plus country_year_product_ext deduplicated to the latest HS revision per country-product, ext restricted to HS4 headings absent from the base. Country names data/parquet/countries.parquet; income groups data/parquet/gmd.parquet; vessel-weighted absolute port latitude data/parquet/portwatch_ports.parquet. Stratified runs: data/parquet/instruments_anti_product_space_strata.parquet.Query
Pairs whose holder sets are furthest apart in income are 4.4 times as likely to be in the deficit tail as pairs whose holders match (4.33% against 0.99%). On port latitude the ratio is 5.4. This is what habitat filtering looks like.
Income rank scores Low income 1, Lower middle 2, Upper middle 3, High income 4, averaged over each heading’s RCA >= 1 holders; the gap is the absolute difference between the two headings’ averages. Latitude is the vessel-weighted absolute port latitude of each heading’s holders, in degrees. Mean income-rank gap: 0.438 over all tested pairs, 0.653 in the deficit tail, 0.172 in the surplus tail. Mean absolute-latitude gap: 7.12 degrees, 10.75 and 3.57. Deciles are of the tested set, so each holds about 9,310 pairs.
CEPII BACI V202501, data/parquet/country_year_product/ and country_year_product_ext/ (export_value in THOUSANDS of USD; every musd column here is millions of USD, thousands / 1000). Primary specification A: HS92 base catalog plus country_year_product_ext deduplicated to the latest HS revision per country-product, ext restricted to HS4 headings absent from the base. Country names data/parquet/countries.parquet; income groups data/parquet/gmd.parquet; vessel-weighted absolute port latitude data/parquet/portwatch_ports.parquet. Covariate gaps: data/parquet/instruments_anti_product_space_gap.parquet.Query
The range runs from -1.262 (Gabon, diversity 18) to +2.086 (Bangladesh, diversity 87). Mean basket tension rises monotonically with income group, from -0.434 in Low income to +0.207 in High income.
Mean over 190 economies is -0.025 in 2024 against -0.261 in 1995. Economies with very small baskets have very few internal pairs, so the positive ranking is restricted to baskets with at least 50 tested pairs; the negative ranking is not restricted, and its entries range from 55 to 561 basket pairs. 4 economies carry no World Bank income group in gmd.parquet and are shown as their own row rather than folded into another. A high tension score is not a diagnosis of a fragile basket: an oil exporter scores negative because oil is unlike everything else, which is a description of oil.
CEPII BACI V202501, data/parquet/country_year_product/ and country_year_product_ext/ (export_value in THOUSANDS of USD; every musd column here is millions of USD, thousands / 1000). Primary specification A: HS92 base catalog plus country_year_product_ext deduplicated to the latest HS revision per country-product, ext restricted to HS4 headings absent from the base. Country names data/parquet/countries.parquet; income groups data/parquet/gmd.parquet; vessel-weighted absolute port latitude data/parquet/portwatch_ports.parquet. Country layer: data/parquet/instruments_anti_product_space_country.parquet.Query
SELECT iso3, diversity, exports_musd, basket_tension, n_pairs_in_basket,
share_negative_pairs, income_group, lat_band FROM read_parquet('data/parquet/instruments_anti_product_space_country.parquet') WHERE year = 2024
ORDER BY basket_tension
For Chinese Taipei (diversity 171, basket tension +0.221 over 465 pairs), the most statistically incompatible target is HS4 2711, Petroleum gases and other gaseous hydrocarbons, whose worst antagonist among the economy’s current specialisations is 8549, Electrical and electronic waste and scrap, at z = -5.33.
One row per (economy, target heading) for every heading with ubiquity at least 25 that the economy does not already hold: 66,798 rows over 190 economies and 432 targets. The antagonist is the economy’s own existing specialisation with the most negative z against the target. 66,792 of 66,798 rows are negative, so a ranking by z is close to a ranking of everything: the density column is the counterweight. Density here is a Hidalgo-Hausmann proximity density recomputed on this same HS4 matrix (proximity = co-occurrence divided by the larger ubiquity, range 0.0013 to 0.5466). It is not the HS6 density number used on /ladder and the two are not interchangeable. This is a description of the co-occurrence record, not a recommendation: a negative z is not a reason for an economy not to enter a heading.
CEPII BACI V202501, data/parquet/country_year_product/ and country_year_product_ext/ (export_value in THOUSANDS of USD; every musd column here is millions of USD, thousands / 1000). Primary specification A: HS92 base catalog plus country_year_product_ext deduplicated to the latest HS revision per country-product, ext restricted to HS4 headings absent from the base. Advisory: data/parquet/instruments_anti_product_space_advisory.parquet (2024).Query
SELECT iso3, target_hs4, target_name, antagonist_hs4, antagonist_name, z,
density, target_ubiquity FROM read_parquet('data/parquet/instruments_anti_product_space_advisory.parquet') WHERE iso3 = 'S19' ORDER BY z LIMIT 12
2,856
13.240
2.429
-5.45
100
1,568
2,699
13.270
2.518
-5.27
200
1,397
2,622
13.160
2.572
-5.12
400
1,325
2,525
13.227
2.509
-5.27
700
1,288
2,515
13.163
2.470
-5.33
1,000
1,285
2,500
13.140
2.484
-5.29
The deficit tail is a decreasing function of the replicate count and is still moving at 200: 2,536 pairs at 25 replicates, 1,397 at 200, 1,285 at 1,000. At the panel setting the null standard deviation is under-estimated and the tail is about 8.7% too large.
The honest headline is therefore about 1,300 pairs, not 1,397. The panel is computed at 200 replicates for all 30 years so the years stay comparable, and that decision costs this bias, which is why it is stated wherever the count appears. The headline pair’s z is insensitive to the same move: -5.27 at 100, -5.12 at 200, -5.29 at 1,000.
Replicate sweep: data/parquet/instruments_anti_product_space_sweep_replicates.parquet (2024, single chain, seed 20260802).Query
SELECT value, n_sig_exclusion, n_sig_attraction, headline_exp_2709x6206,
headline_sd_2709x6206, headline_z_2709x6206 FROM read_parquet('data/parquet/instruments_anti_product_space_sweep_replicates.parquet') ORDER BY value
Query
SELECT value, n_sig_exclusion, n_sig_attraction, share_sig_exclusion,
headline_exp_2709x6206, headline_sd_2709x6206, headline_z_2709x6206 FROM read_parquet('data/parquet/instruments_anti_product_space_sweep_seed.parquet')
ORDER BY value
20
+8.60
At a threshold of 0.50 the headline pair has 3 joint holders, not zero. From 0.75 upward it is empty at every threshold tested, and the z stays negative throughout, from -5.91 at 0.75 to -2.40 at 3.0, where only 1 pair in the whole matrix survives BH.
The tail size falls steeply with the threshold because a stricter RCA rule thins the matrix and the tested set: at 3.0 only 1,081 pairs clear the ubiquity floor at all. The direction of the headline pair is robust to the threshold; the exact zero is not, and no sentence on this page should be read as if it were.
SELECT value, n_countries, n_links, n_pairs_tested, n_sig_exclusion,
obs_2709x6206, exp_2709x6206, z_2709x6206, obs_6110x6203, z_6110x6203 FROM read_parquet('data/parquet/instruments_anti_product_space_sweep_rca.parquet') ORDER BY value
n/a
none
no
The floor decides the size of the tested set, from 760,761 pairs at a floor of 5 to 0 at a floor of 100. The deficit tail peaks near the setting used here and the share of tested pairs in it rises monotonically, because a higher floor keeps only the pairs with enough mass to reject.
At a floor of 40 or above the headline pair is no longer testable at all: its smaller ubiquity is 37. That is a limit of the design, not a result. The floor of 25 used throughout is a power choice: below it, the null standard deviation for a rare pair is too small for a normal approximation to be trusted.
SELECT value, n_pairs_tested, n_sig_exclusion, n_sig_attraction,
share_sig_exclusion, bh_p_cutoff, headline_tested_2709x6206 FROM read_parquet('data/parquet/instruments_anti_product_space_sweep_ubiquity.parquet') ORDER BY value
Query
SELECT value, n_countries, n_links, n_pairs_tested, n_sig_exclusion,
obs_2709x6206, exp_2709x6206, z_2709x6206, obs_6110x6203, z_6110x6203 FROM read_parquet('data/parquet/instruments_anti_product_space_sweep_sizefloor.parquet') ORDER BY value
CEPII BACI V202501, data/parquet/country_year_product/ and country_year_product_ext/ (export_value in THOUSANDS of USD; every musd column here is millions of USD, thousands / 1000). Primary specification A: HS92 base catalog plus country_year_product_ext deduplicated to the latest HS revision per country-product, ext restricted to HS4 headings absent from the base. Specification sweep: data/parquet/instruments_anti_product_space_sweep_spec.parquet (2024).
Query
SELECT value, description, n_countries, n_headings, n_links, world_exports_musd,
n_pairs_tested, n_sig_exclusion, obs_2709x6206, exp_2709x6206, z_2709x6206, u_2709, u_6206,
obs_6110x6203 FROM read_parquet('data/parquet/instruments_anti_product_space_sweep_spec.parquet') ORDER BY value