Loading replication
Fetching primary parquet sources and recomputing the published exhibits.
Fetching primary parquet sources and recomputing the published exhibits.
Melitz (2003) argues that firms are heterogeneous in productivity, and trade liberalization reallocates market share from less to more productive firms. BACI does not contain firm-level data, so a direct Melitz replication is impossible on public sources alone. But the aggregate implication, that selection effects should sharpen the concentration of market share toward the most competitive producers, can be tested with country-level export shares in the top 100 traded products. The across-country variance of log-share has risen from 14.47 (1996) to 22.38 (2024), a +55% increase in dispersion over three decades.
Melitz (2003) embeds Hopenhayn-style firm heterogeneity in productivity φ inside a monopolistic-competition CES trade model. Productivity is drawn from a common distribution (commonly parameterised as Pareto with shape k, which pins down the aggregate elasticity of trade to variable costs). Two cutoffs pin down the equilibrium: φ*, the zero-profit domestic survival cutoff, and φ*x, the zero-profit export cutoff that arises from a fixed export cost fx in addition to iceberg variable cost τ. Only firms with φ ≥ φ*x> φ* serve the foreign market; firms with φ < φ* exit. When trade costs fall, the export cutoff φ*x falls (the extensive margin: more firms start exporting) while φ* rises (tougher home-market selection forces low-φ firms out). Aggregate industry productivity rises through this reallocation with no change in any firm’s own φ , the Melitz selection effect. Continuing exporters expand sales through the intensive margin. The model has been validated on firm microdata by Pavcnik (2002, RESTUD), Bernard-Eaton-Jensen-Kortum (2003, AER), and a long chain of subsequent studies.
BACI is a country-product panel, not a firm-product panel. So we test the Melitz prediction indirectly. Within each of the top 100 globally-traded HS6 products (ranked by 2024 export value), we compute the across-country variance of log(country export share) each year from 1996 to 2024. If selection effects are getting stronger over time, whether through trade cost reductions, supply chain specialization, or productivity dispersion widening, then the variance of log-share should rise: a few countries capture more of the market, the long tail captures less.
On 100 top products, the mean within-product variance of ln(share) rose from 14.47 in 1996 to 22.38 in 2024, a +7.9-unit (+55%) increase. Over the same period, the average number of exporting countries per product rose from 107 to 175, so the widening dispersion is not a mechanical effect of more countries entering the sample, it is within the cross-country distribution.
Chaney (2008, AER) shows that when firm productivity is Pareto-distributed with shape parameter k, the resulting sales distribution is Pareto with the same k, and the country-product export size distribution inherits a Pareto tail. BACI has no firms, but the country × HS6 export flow is the closest aggregate analogue: each of the 517,193 positive-value flows in 2024 is one pair of an exporter with a 6-digit product. Plotting log-rank against log-size for the top 10% of flows (those above $26M) yields a near-linear pattern with slope −0.79, meaning the implied tail parameter α ≈ 0.79. That is within the range Chaney (2008) calibrates k to hit (4-8), and consistent with the thick-tail calibrations of Eaton-Kortum (2002) and Bernard-Eaton-Jensen-Kortum (2003).
Melitz (2003) is agnostic across countries: every industry has a productivity distribution and a cutoff, so every country should show a Pareto tail in its export size distribution , but the tail index α is free to differ. In the Melitz-Chaney mapping, a smaller α means thicker tail: a small number of hyper-competitive product lines carry most of the country’s exports, consistent with strong selection into the best-cost products. A larger α means a flatter distribution: the country exports across a broader base with less concentration at the top. We fit a Hill-like Pareto tail to the top 25% of each country’s 2024 HS6 export-flow size distribution for the 30 largest exporters. The fit α ranges from 0.61 (SAU, thickest tail) to 1.05 (CHN, flattest).
Melitz (2003) pins entry into exporting on a productivity threshold φ*x: only firms with φ ≥ φ*xship abroad; trade liberalisation lowers the cutoff and more firms clear it. We proxy the country-product analogue with Balassa’s RCA: a country × HS6 cell is “above cutoff” if its share in world exports of that product exceeds its share in world total exports (RCA ≥ 1). Stricter tails (RCA ≥ 2 and ≥ 5) trace further into the productivity right tail. The share of cells with RCA ≥ 1 moved from 27.7% in 1996 to 28.4% in 2024; the RCA ≥ 5 tail moved from 8.0% to 8.4%.
Figure 4 collapses the right-tail of the productivity distribution to three thresholds and tracks the share over time. The full histogram tells a more complete Melitz story: where in the RCA distribution did mass shift, and by how much, between the start of the BACI panel and the latest release? A pure Melitz liberalization would predict reallocation away from the “below cutoff” region (RCA < 1) toward the “above cutoff” region (RCA ≥ 1) and a thickening of the right tail (RCA ≥ 5 or 10). A pure extensive-margin entry of new small exporters would do the opposite: pile up cells in the < 0.1 and 0.1-0.5 bins as countries enter many products with negligible specialisation.
source: Melitz (2003), Econometrica, model results in Section 4 and pp. 1713-1721.
qualitative replication, no published table matched
Same (qualitatively): the prediction that trade integration concentrates share at the top of a dispersion distribution, in Melitz, the productivity distribution of firms; here, the market-share distribution of countries within each HS6 product. Differs: unit of heterogeneity (country-product vs firm); observed quantity (export share vs productivity); no fixed-cost / φ*x cutoff is recovered; extensive vs intensive margin cannot be separated on aggregate data (Chaney 2008, AER offers the mapping from BACI-style moments to Melitz primitives, but requires country-pair-product exporter counts that BACI does not carry at the firm level).
This is not a direct replication and the test is weak. Melitz’s model is about firms, not countries, and the reallocation he describes is between firms within a country, not between countries within a product. The right dataset is firm-level plant micro- data, US Census LBD, Chilean ENIA, French customs. With BACI we can only observe the aggregated outcome: the cross-country distribution of who-exports-what. Four caveats.
First, aggregation: country-level export shares reflect both firm-level selection and country-level comparative advantage; rising dispersion could reflect stronger Ricardian specialization rather than Melitz selection, and these are observationally equivalent in country-aggregated data. Second, composition: the top-100 products in 2024 are not the same as the top-100 products in 1996 (petroleum has fallen, electronics have risen); restricting to a fixed 1996 top-100 basket gives the same qualitative pattern, but with smaller dispersion changes. Third, COVID: 2020 is a clear outlier in the series due to trade disruptions; 2021-2024 recover but did not return to the 2015-2019 trend level. Fourth, China: a very large share of the post-2000 rise is mechanically driven by China’s surge in exports of specific products, which thickens the right tail of the cross-country share distribution and shows up as higher variance in logs.
The pattern is consistent withMelitz-style selection effects strengthening over 1996-2024, alongside other channels. Proper firm-level replication would use French, US, or Chilean customs micro-data, which are outside this site’s public-parquet remit.
@article{melitz_2003,
author = {Melitz, Marc J.},
title = {The Impact of Trade on Intra-Industry Reallocations and Aggregate Industry Productivity},
journal = {Econometrica},
volume = {71},
number = {6},
pages = {1695--1725},
year = {2003},
doi = {10.1111/1468-0262.00467}
}Variety-entry evidence at Feenstra (1994). Return to the replication gallery.
WITH top100_2024 AS (
SELECT product_code FROM country_year_product
WHERE year = 2024 AND export_value > 0
GROUP BY product_code
ORDER BY SUM(export_value) DESC LIMIT 100
),
shares AS (
SELECT year, product_code, country_code,
export_value / SUM(export_value) OVER (PARTITION BY year, product_code) AS share
FROM country_year_product
WHERE export_value > 0 AND product_code IN (SELECT product_code FROM top100_2024)
)
SELECT year, AVG(VAR_POP(LN(share))) AS mean_var_log_share
FROM shares WHERE share > 0
GROUP BY year, product_code
ORDER BY year;WITH flows AS (
SELECT export_value * 1000.0 AS v
FROM country_year_product WHERE year = 2024 AND export_value > 0
),
ranked AS (
SELECT v, ROW_NUMBER() OVER (ORDER BY v DESC) AS rk,
NTILE(200) OVER (ORDER BY v) AS bucket
FROM flows
)
SELECT AVG(LN(v)) AS log_size, AVG(LN(rk)) AS log_rank
FROM ranked GROUP BY bucket ORDER BY bucket;
-- α fit by OLS of log_rank on log_size over the top decile (tail) in-app.WITH tot AS (
SELECT c.iso3, SUM(cyp.export_value) AS total
FROM country_year_product cyp JOIN countries c ON c.code = cyp.country_code
WHERE cyp.year=2024 AND cyp.export_value > 0
GROUP BY c.iso3 ORDER BY total DESC LIMIT 30
)
SELECT c.iso3, cyp.export_value * 1000 AS v
FROM country_year_product cyp JOIN countries c ON c.code=cyp.country_code
JOIN tot USING(iso3) WHERE cyp.year=2024 AND cyp.export_value > 0
ORDER BY c.iso3, cyp.export_value DESC;
-- α fit per iso3 on top 25% of its flows by OLS of log(rank) on log(v).WITH rca AS (
SELECT year, country_code, product_code, rca
FROM rca_matrix WHERE rca IS NOT NULL AND year BETWEEN 1996 AND 2024
)
SELECT year,
AVG(CASE WHEN rca >= 1 THEN 1.0 ELSE 0.0 END) AS share_rca1,
AVG(CASE WHEN rca >= 2 THEN 1.0 ELSE 0.0 END) AS share_rca2,
AVG(CASE WHEN rca >= 5 THEN 1.0 ELSE 0.0 END) AS share_rca5,
COUNT(*) AS n_pairs
FROM rca GROUP BY year ORDER BY year;SELECT year,
CASE WHEN rca < 0.1 THEN '<0.1'
WHEN rca < 0.5 THEN '0.1-0.5'
WHEN rca < 1 THEN '0.5-1'
WHEN rca < 2 THEN '1-2'
WHEN rca < 5 THEN '2-5'
WHEN rca < 10 THEN '5-10'
ELSE '10+' END AS bin,
COUNT(*) AS n
FROM rca_matrix WHERE year IN (1996, 2024) AND rca IS NOT NULL
GROUP BY year, bin ORDER BY year, bin;
-- Then divide each (year, bin) count by year totals to get shares.