Personal working proposal · evidence checked 10 July 2026

Research training outside the university. Is it worth testing?

A small, serious program built around supervised research, close feedback, and judgment: choosing good questions, making defensible inferences, and using AI without outsourcing the thinking.

This page is meant to be marked up. Select a passage, or open the Hypothes.is panel at the right edge, to annotate it. Corrections and skeptical reactions are especially useful.

The short version

§1 The offer

Content is not the scarce part. Observation, correction, and trust may be.

What participants would actually get

The central product is supervised practice in making research decisions, not a compressed degree.

Foundation

Shared language and diagnostic work

Short, hands-on sessions in statistical reasoning, causal claims, modeling under uncertainty, and reproducible workflows. Enough unassisted work to reveal what a participant understands before AI is added.

Project

A real question with visible decisions

Each participant develops a small research project. Mentors see the discarded ideas, revisions, checks, and mistakes, not only the final document. AI use is logged and audited rather than hidden or prohibited.

Feedback

A conference organized around reading

Work is circulated in advance. Discussants identify the strongest claim, the weakest assumption, and the next useful test. The aim is detailed criticism and revision, not a sequence of polished slide talks.

Signal

Evidence of how someone thinks

Participants leave with an audited project, a feedback record, and, where warranted, a specific letter from someone who observed their reasoning. There is no degree or accreditation.

§2 Evidence

Published findings, early labor-market evidence, and my own hypotheses should not be presented as one category.

What changed, and how sure am I?

Published RCT

Carefully designed AI tutoring can teach some content well. In a crossover trial with 194 Harvard physics students, a structured AI tutor produced larger immediate learning gains in less time than an active-learning class. It used expert-written materials and prompts; the authors do not claim this will generalize to higher-order synthesis or every subject.

Published field experiment

Unguided AI assistance can improve practice performance without improving learning. In a study of nearly 1,000 high-school mathematics students in Turkey, unrestricted GPT access helped with practice but led to worse performance once access was removed. A constrained tutor avoided the measured harm. This is one subject and setting, but it supports deliberate AI-free checks.

Working paper

There is a warning sign for entry-level, AI-exposed work. A Stanford Digital Economy Lab paper using US payroll data reports a 16% relative employment decline for 22- to 25-year-olds in the most AI-exposed occupations, after controls for firm-level shocks. The pattern is consistent with AI displacement; it is not a settled causal estimate and sits within a broader hiring slowdown.

Design hypothesis

Polished output is becoming a weaker signal of independent skill. This seems likely when competent code and prose are cheap to generate, but I do not have direct evidence that employers or PhD admissions now put more weight on observed process and detailed letters.

Design hypothesis

Judgment is the scarce skill this program can improve. Question choice, identification, model criticism, and knowing what not to delegate are plausible targets. Whether a short program can teach and credibly assess them is exactly what a pilot should test.

Sources, effect sizes, and limits
  • Kestin et al. (Scientific Reports, 2025): Harvard physics crossover RCT, N=194; reported effect sizes of roughly 0.73 to 1.3 standard deviations depending on specification. Outcomes were immediate post-tests in two lessons, not long-term retention or independent research.
  • Bastani et al. (PNAS, 2025): randomized field experiment in one Turkish high school. The result supports guardrails in this context, not a general ban on AI-assisted learning.
  • Brynjolfsson, Chandar, and Chen (2025, revised November 2025): administrative payroll evidence and an event-study design. The authors describe the findings as early evidence consistent with the AI-impact hypothesis.
§3 Curriculum

AI-on work and AI-off checks serve different purposes.

What may still be worth teaching

AI changes the balance, not the need for foundations. Participants should sometimes derive, code, or reason without assistance when doing so builds understanding or makes competence observable. Once the idea is understood, routine execution can be delegated and checked.

Questions

Problem choice and research design

Turning a broad concern into a tractable question; distinguishing prediction, description, and causal inference; deciding what evidence would change the conclusion.

Statistics

Meaning before machinery

What an estimate and its uncertainty mean, what assumptions identify a claim, which comparisons are credible, and how selection or measurement can break the analysis.

Modeling

Explicit uncertainty

Fermi estimation, calibration, forecasting, sensitivity analysis, and cost-effectiveness models whose parameters can be inspected and challenged.

AI workflow

Delegate, verify, document

Using agents for search, code, data work, and critique; checking provenance and calculations; constructing adversarial tests; keeping an audit trail; and protecting confidential data.

Communication

Make the claim inspectable

Writing short, calibrated conclusions; linking claims to evidence; separating results from interpretation; and responding constructively to strong criticism.

Plain-language glossary

Causal identification: the argument that a comparison isolates the effect of interest rather than some other difference. Calibration: making confidence match accuracy over repeated judgments. Sensitivity analysis: checking whether a conclusion survives plausible changes in assumptions. Reproducibility: making it possible for someone else to trace and rerun the analysis. Audit trail: a record of sources, decisions, tool use, and checks.

§4 Landscape

The gap is a combination of features, not an empty field.

What already exists

There are strong adjacent programs. The AEA Summer Program provides intensive economics preparation and credit. SICSS combines computational-social-science instruction with collaborative projects. PREDOC now offers asynchronous research-tooling courses, including agentic AI. GovAI and several AI-safety fellowships provide close mentorship in narrower domains.

So the claim is not "nobody does research training." The possible gap is this particular bundle: broad quantitative social science, deliberate AI-on and AI-off work, continuous observation of reasoning, and a closing conference centered on pre-read criticism. I have not found a close match. That may indicate an opportunity, or simply weak demand.

Comparison with selected current programs
ProgramMain offerWhere it overlapsDifference from this sketch
AEA Summer ProgramTwo-month residential economics preparation with coursework and creditTechnical training, mentoring, research methodsEconomics and PhD pipeline; more course-centered and US-focused
SICSSOne- or two-week computational-social-science institutes, often with group projectsMethods, open science, projects, peer communityShorter project cycle; usually field- and method-specific
PREDOC coursesAsynchronous tools for research assistants, including AI workflowsPractical research tooling and AIOnline content rather than an observed project and feedback conference
GovAI Research ScholarsMentored research in AI governanceClose mentorship and real outputAdvanced and domain-specific rather than broad entry-stage training
§5 The pilot

Build the smallest version that can prove the premise wrong.

A first test, not an institute

A useful pilot might be one intensive week, a six- to eight-week part-time project, and one feedback day, with perhaps 12 to 16 participants. Those numbers are illustrative; mentor capacity should determine the scale.

Before recruitment, the organizers should state what would count as success and what would make them stop. The primary question is whether participants make better research decisions on a new, unassisted task, not whether they enjoy the program or produce polished documents.

What to measure

  • Baseline and endline reasoning on unseen AI-free tasks
  • Blind expert ratings of question choice, design, and calibration
  • Whether project claims survive replication and adversarial checks
  • Mentor hours, attrition, and participant opportunity cost
  • Usefulness of the resulting letters to real selectors

Reasons to stop or redesign

  • No improvement on new tasks beyond selection and practice effects
  • Mentor time per participant is too high to repeat
  • Equivalent feedback is already available more cheaply
  • AI logs reward compliance rather than honest experimentation
  • The signal is not trusted by employers or graduate programs

The comparison matters. The pilot should be judged against realistic alternatives: self-study with an AI tutor, a standard methods course, research-assistant work, or direct mentorship. A pre/post improvement by itself would not establish value.

§6 People and access

Who might benefit

  • People considering a research career who need a realistic trial before committing to a degree or job path.
  • Domain specialists who want quantitative tools and can bring substantive knowledge to a model or empirical project.
  • Quantitative researchers moving toward higher-stakes questions in policy, global priorities, or the social and economic effects of AI.
  • Researchers who mainly need the feedback stage and would bring an existing paper for intensive pre-read criticism.

Selection and access are unresolved. A paid residential model would exclude many of the people who might benefit most. Any pilot should publish its selection rubric, offer remote or funded participation where feasible, and distinguish demonstrated reasoning from institutional pedigree. These are design requirements, not promises I can yet make.

§7 Provenance

Where this comes from

I first sketched a broader summer institute around 2020. Since then I have run smaller modeling workshops and supervised early-career researchers. The parts that still seem distinctive are the observed project, explicit uncertainty, and feedback conference. The AI-era framing is newer and more tentative.

I co-direct The Unjournal, which commissions public expert evaluation of research. There is an overlap in emphasis on careful reading and structured criticism, but this is a separate personal exploration and does not speak for The Unjournal.

The ask

Tell me which premise fails.

I am collecting evidence before deciding whether to organize a pilot. Reactions are useful from prospective participants, mentors, selectors, funders, and people who think an existing program already does this better.

How you are responding (optional)
or email me