AI-integrated research; a novel tradeoff and partial solution (part 1 of n)
Moving all your project ideas forward by "one step" but also adding a new bottleneck
File Under: Integrating AI into research
In reaction to scott cunningham’s very interesting new post about hitting some walls with integrating Claude into research:
My initial reaction is a bit different. A key theme here: AI doesn’t just make research faster — it reallocates bottlenecks.
I just got back yesterday from the NBER Economics of Health Spring meeting, where I was observing, not presenting. I’ll have more to say later on an interesting paper on NAFTA, where two teams have asked basically the same question at around the same time—more on that later.
And I saw several (especially junior scholars) folks taking notes on the presentations. And I recall many years ago doing the same thing—notebooks instead of laptop/notepad. I’ve chatted with Joe Price about this recently—and he continues to have (likely 1000s) of notes from presentations, meetings, and so on. I’ve mostly stopped taking notes and just relying on the idea that “if it is important enough, I’ll remember it at some point”. This is a reaction to having pages of notes (I would often scan them and add them into folders, transcribe them, and so on) that just accumulate and rarely turn into a project.
So the bottleneck here was a combination of time and ranking projects on promise. That is, we basically have to pay a fixed cost to set up a project (gather data, data management/cleaning…) just to try to figure out whether it held any promise that we should then invest further in.
The AI-induced collapse of the feasibility filter.
Claude Code (and similar) lets us move many projects one step forward. The fixed cost for set up is getting very close to zero. Now we can assess projects for their promise and rank them with much more information than before. But we don’t actually have much more time to pursue and implement(i.e. finish) more projects (i.e. moving from feasible to completed). We just now know which ones are likely to work out—we can move them onto a separate notebook and include initial results, but the number of projects on this (now much smaller) notebook is still orders of magnitude more than the number of projects we can finish.
And, as Scott says, we also now need to invest time in human verification of AI results, which costs us time compared to older versions of project development. This is a new-ish bottleneck in our projects. It’s not completely new because we always needed to review the work of our co-authors or (especially) student/RAs before moving projects forward. But, the type of errors from AI are new and come in unexpected places.
The A and B Teams
One idea here that I’m kicking around is to redeploy typical RA tasks to help ease this new bottleneck. The idea would be to stop training RAs as co-authors (unfortunately this is bad for them) and basically train them as the “B Team”. Your AI (A Team) has produced an answer that is probably right—assuming you’ve done some due diligence and light kicking of the tires— but has a chance of having problems. You bring in the B Team to start the analysis from scratch and see if the answers line up. This also allows your final project to be “human only” and not raise any flags at journals, etc. The big win here is that there will basically be no projects that you assign an RA that fail. They all have had a first look (and maybe several) by AI, so you will pay no more fixed costs to set up projects and wait for an RA to work on them for months only to find nothing (or run into a data or other problem). And while the Team B is ‘confirming’ Team A’s result, you can write the paper or temporarily move on to another project and circle back when Team B is done with that task and moves on to the next project immediately. The idea is that Team B is not a human verifier in the sense of reviewing AI code but instead is a “normal RA” but with very fluffy AI generated guard rails:
You already basically know the answer and will assume (at least in the beginning) that any differences found by Team B are not correct. You might even tell the RA—run the analysis, the results are likely in the ball park of X. If you don’t get X, very carefully review and troubleshoot. And do not use AI.






My guess is that RAs will be willing to abide by clear rules to not use AI for a specific project. One idea is to tell them that the two of you will go over their code line-by-line and they will explain their thinking, etc. This is akin to the move back to oral exams and in-class exams as universities shift away from take-home assessments. Not perfect, but perhaps close enough.
Good thinking about different bottlenecks! But can we make sure that our RAs don't use AI? This is probably a lost battle. The best thing we can hope for is that they will use different AI models somewhat differently.