Personal working proposal · evidence checked 10 July 2026
Research training outside the university. Is it worth testing?
A small, serious program built around supervised research, close feedback, and judgment: choosing good questions, making defensible inferences, and using AI without outsourcing the thinking.
This page is meant to be marked up. Select a passage, or open the Hypothes.is panel at the right edge, to annotate it. Corrections and skeptical reactions are especially useful.
The short version
- The program would not mainly sell lectures or coding instruction. Those are increasingly cheap and available.
- It would provide scarce inputs: an observed research process, demanding supervision, an audited project, and credible feedback.
- The first step should be a small pilot with pre-declared tests, not a full institute with a name, venue, and permanent staff.
- There are no dates or applications. I am trying to learn whether the combination is useful and sufficiently distinct from existing programs.
Content is not the scarce part. Observation, correction, and trust may be.
What participants would actually get
The central product is supervised practice in making research decisions, not a compressed degree.
Shared language and diagnostic work
Short, hands-on sessions in statistical reasoning, causal claims, modeling under uncertainty, and reproducible workflows. Enough unassisted work to reveal what a participant understands before AI is added.
A real question with visible decisions
Each participant develops a small research project. Mentors see the discarded ideas, revisions, checks, and mistakes, not only the final document. AI use is logged and audited rather than hidden or prohibited.
A conference organized around reading
Work is circulated in advance. Discussants identify the strongest claim, the weakest assumption, and the next useful test. The aim is detailed criticism and revision, not a sequence of polished slide talks.
Evidence of how someone thinks
Participants leave with an audited project, a feedback record, and, where warranted, a specific letter from someone who observed their reasoning. There is no degree or accreditation.
Published findings, early labor-market evidence, and my own hypotheses should not be presented as one category.
What changed, and how sure am I?
Carefully designed AI tutoring can teach some content well. In a crossover trial with 194 Harvard physics students, a structured AI tutor produced larger immediate learning gains in less time than an active-learning class. It used expert-written materials and prompts; the authors do not claim this will generalize to higher-order synthesis or every subject.
Unguided AI assistance can improve practice performance without improving learning. In a study of nearly 1,000 high-school mathematics students in Turkey, unrestricted GPT access helped with practice but led to worse performance once access was removed. A constrained tutor avoided the measured harm. This is one subject and setting, but it supports deliberate AI-free checks.
There is a warning sign for entry-level, AI-exposed work. A Stanford Digital Economy Lab paper using US payroll data reports a 16% relative employment decline for 22- to 25-year-olds in the most AI-exposed occupations, after controls for firm-level shocks. The pattern is consistent with AI displacement; it is not a settled causal estimate and sits within a broader hiring slowdown.
Polished output is becoming a weaker signal of independent skill. This seems likely when competent code and prose are cheap to generate, but I do not have direct evidence that employers or PhD admissions now put more weight on observed process and detailed letters.
Judgment is the scarce skill this program can improve. Question choice, identification, model criticism, and knowing what not to delegate are plausible targets. Whether a short program can teach and credibly assess them is exactly what a pilot should test.
Sources, effect sizes, and limits
- Kestin et al. (Scientific Reports, 2025): Harvard physics crossover RCT, N=194; reported effect sizes of roughly 0.73 to 1.3 standard deviations depending on specification. Outcomes were immediate post-tests in two lessons, not long-term retention or independent research.
- Bastani et al. (PNAS, 2025): randomized field experiment in one Turkish high school. The result supports guardrails in this context, not a general ban on AI-assisted learning.
- Brynjolfsson, Chandar, and Chen (2025, revised November 2025): administrative payroll evidence and an event-study design. The authors describe the findings as early evidence consistent with the AI-impact hypothesis.
AI-on work and AI-off checks serve different purposes.
What may still be worth teaching
AI changes the balance, not the need for foundations. Participants should sometimes derive, code, or reason without assistance when doing so builds understanding or makes competence observable. Once the idea is understood, routine execution can be delegated and checked.
Problem choice and research design
Turning a broad concern into a tractable question; distinguishing prediction, description, and causal inference; deciding what evidence would change the conclusion.
Meaning before machinery
What an estimate and its uncertainty mean, what assumptions identify a claim, which comparisons are credible, and how selection or measurement can break the analysis.
Explicit uncertainty
Fermi estimation, calibration, forecasting, sensitivity analysis, and cost-effectiveness models whose parameters can be inspected and challenged.
Delegate, verify, document
Using agents for search, code, data work, and critique; checking provenance and calculations; constructing adversarial tests; keeping an audit trail; and protecting confidential data.
Make the claim inspectable
Writing short, calibrated conclusions; linking claims to evidence; separating results from interpretation; and responding constructively to strong criticism.
Plain-language glossary
Causal identification: the argument that a comparison isolates the effect of interest rather than some other difference. Calibration: making confidence match accuracy over repeated judgments. Sensitivity analysis: checking whether a conclusion survives plausible changes in assumptions. Reproducibility: making it possible for someone else to trace and rerun the analysis. Audit trail: a record of sources, decisions, tool use, and checks.
The gap is a combination of features, not an empty field.
What already exists
There are strong adjacent programs. The AEA Summer Program provides intensive economics preparation and credit. SICSS combines computational-social-science instruction with collaborative projects. PREDOC now offers asynchronous research-tooling courses, including agentic AI. GovAI and several AI-safety fellowships provide close mentorship in narrower domains.
So the claim is not "nobody does research training." The possible gap is this particular bundle: broad quantitative social science, deliberate AI-on and AI-off work, continuous observation of reasoning, and a closing conference centered on pre-read criticism. I have not found a close match. That may indicate an opportunity, or simply weak demand.
Comparison with selected current programs
| Program | Main offer | Where it overlaps | Difference from this sketch |
|---|---|---|---|
| AEA Summer Program | Two-month residential economics preparation with coursework and credit | Technical training, mentoring, research methods | Economics and PhD pipeline; more course-centered and US-focused |
| SICSS | One- or two-week computational-social-science institutes, often with group projects | Methods, open science, projects, peer community | Shorter project cycle; usually field- and method-specific |
| PREDOC courses | Asynchronous tools for research assistants, including AI workflows | Practical research tooling and AI | Online content rather than an observed project and feedback conference |
| GovAI Research Scholars | Mentored research in AI governance | Close mentorship and real output | Advanced and domain-specific rather than broad entry-stage training |
Build the smallest version that can prove the premise wrong.
A first test, not an institute
A useful pilot might be one intensive week, a six- to eight-week part-time project, and one feedback day, with perhaps 12 to 16 participants. Those numbers are illustrative; mentor capacity should determine the scale.
Before recruitment, the organizers should state what would count as success and what would make them stop. The primary question is whether participants make better research decisions on a new, unassisted task, not whether they enjoy the program or produce polished documents.
What to measure
- Baseline and endline reasoning on unseen AI-free tasks
- Blind expert ratings of question choice, design, and calibration
- Whether project claims survive replication and adversarial checks
- Mentor hours, attrition, and participant opportunity cost
- Usefulness of the resulting letters to real selectors
Reasons to stop or redesign
- No improvement on new tasks beyond selection and practice effects
- Mentor time per participant is too high to repeat
- Equivalent feedback is already available more cheaply
- AI logs reward compliance rather than honest experimentation
- The signal is not trusted by employers or graduate programs
The comparison matters. The pilot should be judged against realistic alternatives: self-study with an AI tutor, a standard methods course, research-assistant work, or direct mentorship. A pre/post improvement by itself would not establish value.
Who might benefit
- People considering a research career who need a realistic trial before committing to a degree or job path.
- Domain specialists who want quantitative tools and can bring substantive knowledge to a model or empirical project.
- Quantitative researchers moving toward higher-stakes questions in policy, global priorities, or the social and economic effects of AI.
- Researchers who mainly need the feedback stage and would bring an existing paper for intensive pre-read criticism.
Selection and access are unresolved. A paid residential model would exclude many of the people who might benefit most. Any pilot should publish its selection rubric, offer remote or funded participation where feasible, and distinguish demonstrated reasoning from institutional pedigree. These are design requirements, not promises I can yet make.
Where this comes from
I first sketched a broader summer institute around 2020. Since then I have run smaller modeling workshops and supervised early-career researchers. The parts that still seem distinctive are the observed project, explicit uncertainty, and feedback conference. The AI-era framing is newer and more tentative.
I co-direct The Unjournal, which commissions public expert evaluation of research. There is an overlap in emphasis on careful reading and structured criticism, but this is a separate personal exploration and does not speak for The Unjournal.
The ask
Tell me which premise fails.
I am collecting evidence before deciding whether to organize a pilot. Reactions are useful from prospective participants, mentors, selectors, funders, and people who think an existing program already does this better.