How to Write a Top 5 Paper: A Case Study
Two teams, nearly the same question, nearly the same data—and a revealing difference in what the papers ask the evidence to do
You almost never get a clean comparison of how papers get framed for top journals. Advice usually arrives after the fact. A paper has already succeeded, and then people reconstruct a story about the size of the question, the novelty of the design, or the brilliance of the execution.
What is rarer is something closer to a natural experiment: two independent teams, working at almost the same time, on almost the same question, with closely related data, reaching nearly the same first-order conclusion—and then positioning the papers very differently.
In February 2026, I found myself in exactly that situation.
---
Two NAFTA/Mortality papers arrive at nearly the same time
Hamid Noghanibehambari and I had been working for close to two years on a paper about NAFTA and mortality. We first sent it out in October 2024. By December 2025 it was at its third journal as a revise and resubmit. We sent the revision back in February 2026.
That same month I was at the NBER Spring Health Meeting in Chicago. When the program came out, I saw that Amy Finkelstein, Matt Notowidigdo, and Steven Shi had a paper on NAFTA and mortality. Around the same time, Hamid and I sent our paper—“The Silk Road of Ashes: Exposure to NAFTA and Adult Mortality”—to the NBER working paper series. It came out the Monday before the conference.
The timing was close enough that Matt emailed to acknowledge he had seen our paper and that their team would be presenting later that week. Their paper—“Trading Goods for Lives: NAFTA’s Mortality Impacts and Implications”—circulated soon afterward. It was later featured in the New York Times a couple weeks later. As far as I know, our paper was not mentioned in that coverage.
Prestige is part of this story. It would be naive to pretend otherwise. Author reputations affect which papers land on programs, which working papers journalists notice, and how readers approach a new result.
But prestige is not the whole story. If you read the two papers side by side, you also see a substantive difference in what the authors ask the evidence to accomplish—and I think this gives a good view about elements of “how to write a top 5”.
---
The same first-order finding
At a broad level, the papers find the same thing. Places more exposed to NAFTA experienced an increase in mortality.
There are differences in design—commuting zones versus PUMAs, different exposure measures, different age ranges, different outcome constructions. Our paper focuses on working-age adults and documents pathways through employment, income, housing wealth, disability, private and public insurance, and transfers. Their paper studies all-age, age-adjusted mortality, emphasizes especially large effects among working-age men, and adds evidence on smoking, self-reported health, and related outcomes.
Those differences matter for interpretation. Some of them create mostly minor disagreements that deserve more attention than they have gotten. But they do not, by themselves, explain a huge difference in a paper’s general-interest fit. The common reduced-form result is too similar.
This is what makes the comparison useful. The case study is not that one team found the effect and the other did not. The case study is what happens after both teams find the effect.
---
What is the paper actually about?
Our paper is fundamentally about NAFTA. It asks what NAFTA did to adult mortality and then investigates a large number of plausible pathways.
The Finkelstein paper is only partly about NAFTA. It uses NAFTA as the central case in a broader question: why do plant closings and trade shocks tend to increase mortality while recessions often reduce it? Put differently, why do different kinds of local economic contractions appear to have opposite effects on health?
That question survives even if NAFTA disappears. One of the clearest tests in the ChatGPT comparison (I had ChatGPT to a deep dive on both papers, including five reviewer agents) was basically: delete the word “NAFTA” from each paper. Most of our empirical object disappears. Their paper still has a puzzle, a cross-shock design, a proposed explanation, and a broader mortality framework.
Their proposed answer is that the sectoral composition of job loss matters. Trade shocks disproportionately reduce manufacturing employment. Typical recessions produce more non-manufacturing employment loss. The paper argues that manufacturing declines raise mortality, while non-manufacturing declines can reduce it. It then uses evidence from NAFTA, the China shock, the Great Recession, and ordinary local employment variation to build a more portable claim.
I do not necessarily find the manufacturing-versus-non-manufacturing distinction fully compelling as a final explanation—either in terms of the results or the strategy. but to each their own.
But that is not the central point of this case study. The point is that the authors locate a puzzle in the literature and make the NAFTA estimate serve as one step in an attempt to resolve it. The claim is more ambitious, more ‘general’, and therefore more interesting to readers who do not primarily care about NAFTA.
---
The welfare calculation
The second major difference is the welfare calculation.
Our paper describes mortality as a hidden cost of trade adjustment and discusses the inadequacy of the safety net. We do not formally put the mortality effect inside a standard welfare calculation. Their paper does. It converts the mortality estimates into changes in life expectancy, puts those changes into a stylized utility framework, and asks whether the implied losses could overturn conventional estimates of NAFTA’s gains through wages and prices. If you wanted a “checklist” for if your paper “looks” like a top five, this is one of the boxes to check.
The exercise is not definitive. It requires extrapolation from local estimates, values for a statistical life-year, and so on. The authors acknowledge these caveats. But the calculation changes the paper’s main point from “NAFTA had another important cost” to “including health may reverse the sign of the canonical welfare evaluation of a major trade agreement.”
That is a much bigger step towards making a paper “general interest”. It brings trade economists, public economists, labor economists, and health economists into the same conversation. It also produces a sentence that is easy to remember and easy to transmit.
Top-five papers often do this. They do not merely document a consequence. They force the consequence to interact with an object that a large literature already treats as central.
---
Boldness is an input into the paper
It is tempting to call all of this framing or marketing—as though the empirical work is the real paper and the broader argument is packaging added afterward. I think that is a mistake.
Finding the larger puzzle is intellectual work. Deciding which pieces of evidence belong in the main argument is intellectual work. Translating a result into a welfare object is intellectual work. So is making a claim general enough that another field can use it. I put a lot of weight on that investment: expanding the possible scope of interest in a given set of results. And often, I find this work to be quite difficult.
ChatGPT suggests the following contrast
—Finkelstein paper is organized hierarchically. It begins with a contradiction: some contractions raise mortality and others lower it. It establishes the NAFTA fact. It proposes a sectoral explanation. It compares the explanation across shocks. It then carries the mortality effect into welfare analysis. Each section tries to raise the stakes of the previous section.
—Our paper is more horizontally additive. We estimate the mortality effect and then examine employment, income, industries, wealth, homeownership, disability, insurance, transfers, and heterogeneity. In some respects, our mechanism evidence is richer and more directly policy relevant. But the abundance of outcomes does not automatically produce a larger contribution--we get tagged as “just switching around the Y variable”. Without a single adjudicating idea, additional results can make a paper more comprehensive without making it more general.
This helps explain a recurring feature of top-five abstracts and introductions. They are usually extremely bold—sometimes almost outlandishly so—about the problem being solved and the importance of the answer. The better papers earn that language. Other papers stretch--everyone has their own threshold about how much to ‘stretch’ their findings. But the common feature is an enormous investment in explaining why the evidence changes how economists should think about something larger than the setting in which it was estimated.
---
Prestige matters, but it is not the whole explanation
None of this means that journal placement or press coverage is a pure meritocracy. It is not. Established authors receive more attention, more favorable priors, and more opportunities to present early work. A manuscript with the same words and different names gets different treatment.
But the framing difference also makes the Finkelstein paper easier to select, describe, cite, and publicize. “NAFTA increased working-age mortality” is an important result. “The mortality effect of an economic contraction depends on which kinds of jobs disappear” is a claim that can organize a much larger conversation. “Health costs may overturn the conventional welfare gains from trade” is a headline.
---
The uncomfortable lesson for careful empirical researchers
Many empirical researchers are trained, correctly, to be cautious. We worry about identification, external validity, multiple mechanisms, and whether a coefficient supports the sentence placed above it. That discipline is essential.
But caution can also become an hurdle to not perform the last stage of intellectual work. We can devote lots of time to making an estimate more credible and comparatively little time to asking what broad question the estimate can answer; Add outcomes but not extend the reach of the analysis; Be cautious about generalization as overclaiming rather than as devote to time to really try to integrate the results with other ideas/fields.
Chatgpt says: The lesson is not to make claims that the evidence cannot bear. The lesson is that evidence rarely announces its own generality. Authors have to construct the bridge from the estimate to the larger question. A top-five paper often has a longer, bolder bridge.
---
What this suggests for our paper--which someone else can do
ChatGPT’s recommendation was that Hamid and I should make the disagreements between the two papers central to our revised paper. No thanks--our paper is now accepted at Canadian Journal of Economics and we’ll move to other things--but someone else could do this and probably get a publication.
The papers agree on the basic sign, but they disagree on important dimensions. Our effects are concentrated among adults ages 25 to 55; theirs appear across age groups. We find a relatively selective cause-of-death pattern; they find increases across a much broader set of causes. We find declines in private insurance and movement toward public coverage; their insurance results differ. The magnitudes also depend on geography, exposure scaling, age adjustment, sample years, cause coding, and the choice of denominator. A harmonized analysis that holds commuting zones, exposure, years, weights, and outcome definitions fixed—and then changes one choice at a time—could find something interesting, who knows..
A remaining question is about publication speed--with our papers now accepted, how big a loss of a “first mover” advantage will their paper have?
---
A practical top-five test--from ChatGPT, I’m still mulling.
This case suggests a set of questions that are useful well before a paper is finished:
1. If the empirical setting disappeared, would the motivating question still matter?
2. Is the main contribution a new fact in one setting, or an explanation that can travel across settings?
3. Can a reader outside the immediate subfield repeat the paper’s central claim in one sentence?
4. Do later sections escalate one argument, or do they mostly add outcomes and robustness checks?
5. Does the evidence change a canonical calculation, resolve a contradiction, or reorganize an established literature?
6. What is the boldest interpretation the evidence can responsibly support—and what additional test would make that interpretation credible?
Two teams can find nearly the same fact. The paper with the higher general-interest ceiling may not have a dramatically more credible estimate. It may instead ask the estimate to do more. That can look like confidence, ambition, or overreach. But it also looks like more/harder work and can be valuable.
---
The two papers
- Amy Finkelstein, Matthew J. Notowidigdo, and Steven Shi, “Trading Goods for Lives: NAFTA’s Mortality Impacts and Implications” (NBER Working Paper 34855, February 2026).
- Hamid Noghanibehambari and Jason Fletcher, “The Silk Road of Ashes: Exposure to NAFTA and Adult Mortality” (NBER Working Paper 34840, February 2026).
Note: The journal-placement premise here is prospective, not a report of completed outcomes. I asked ChatGPT to assume that one paper would land in a top-five economics journal and the other in a strong second-tier general-interest journal, then identify manuscript features that could rationalize that sorting.



A really great post on a topic of great interest. It's something I've ruminated on for the last 30 years--a time during which I've batted .000 on top 5s, so I clearly haven't figured it out!
Really like your post. You’ve coherently written what I’ve fretted about for a long time. Generalizable (large audience) and bold(er) claims seem to be the missing elements for good papers that don’t rise higher. You’re right though that often this is another part of the intellectual work, but I don’t think it gets enough attention in phd training.