Articles
How AI Is Transforming Research: Subjects, Practitioners, Institutions, and Social Impact in the AI Age
Automatically translated from the Japanese original.
1. Introduction
Until a few years ago, I worked as a researcher. With the arrival of general-purpose, high-performance AI such as ChatGPT and Claude, I feel that research has changed dramatically from what it was back then.
Since general-purpose AI like ChatGPT appeared, the research world has been profoundly affected. Submissions to NeurIPS, the premier conference in machine learning, more than doubled in five years, from 9,467 in 2020 to 21,575 in 2025 (NeurIPS, 2025), and Organization Science, a leading management journal, has reported that submissions rose 42% year on year after ChatGPT's public release (Organization Science Editorial Board, 2026).
In this article, I take a fresh look at research as a whole through a comprehensive survey of its object, its actors, its institutions and its social significance, and I discuss what kind of transformation the age of AI demands of it.
2. The Object of Research
2.1 Defining Research
The OECD's Frascati Manual, the international standard for R&D statistics, defines research as "creative and systematic work undertaken in order to increase the stock of knowledge and to devise new applications of that knowledge" (OECD, 2015). The manual also sets out five criteria for judging whether an activity counts as research, giving us a tool for drawing the line that works regardless of discipline.
Novel: it aims to acquire findings that are not yet known
Creative: it rests on original, non-obvious concepts or hypotheses
Uncertain: the outcome (whether it will succeed, and at what cost and over what period) is not determined in advance
Systematic: it is carried out in a planned, budgeted way that leaves records
Transferable / reproducible: it aims to put its results in a form that can be communicated to others and reproduced
One implication of these five criteria is that they also make clear which activities look like research but are not.
Gathering information about things for which answers already exist (this is learning or investigation; it may be new to the person, but it adds nothing to the world's knowledge)
Routine data collection, quality control and market research (systematic, but lacking novelty and creativity)
Routine improvement of existing products (little uncertainty, so it is classified as development)
In short, research can be defined as "adding knowledge that is new to the world, in a verifiable form."
2.2 The Four Acts That Make Up Research
If we break this definition down into actions, research in any field turns out to be a combination of the following four acts.
1. Posing a question
Identifying a question that has no answer yet and is worth answering (it is unresolved, it can in principle be answered, and answering it would change something)
Example: "Can a protein's three-dimensional structure be predicted purely by computation from its amino acid sequence?" This question, unresolved for 50 years, was the starting point for AlphaFold (Jumper et al., 2021)
2. Constructing candidate answers
Building hypotheses, theories, models or artifacts (algorithms or apparatus)
Example: Watson and Crick's wire model of the DNA double helix (Watson & Crick, 1953) was a physical construction of a candidate answer
3. Testing
Establishing how likely the answer is to be correct through evidence, argument or experiment
Falsifiability (Popper, 1959): a claim is testable only if one can say "what observation, if made, would refute this answer"
Example: general relativity produced the falsifiable prediction that "the apparent positions of stars should shift during a solar eclipse," and it survived observation (Dyson, Eddington & Davidson, 1920)
4. Sharing and making verifiable
Publishing results in a form that others can replicate, criticize and reuse
Scope: not only publishing papers, but also releasing data, code and experimental procedures
These four acts form a loop rather than a straight line: a failed test gives rise to a new question, and criticism of shared results gives rise to the next hypothesis.
2.3 Three Types of Research Approach: Classification by Research Purpose
Here I present two classifications that serve different roles.
The classic three-way classification is used in national research statistics, government budgets and corporate R&D accounting; everything is recorded under these categories. However, it implicitly assumes a one-dimensional line running "basic → applied → development," and it has no place for research that pursues understanding and practical use at the same time.
Pasteur's quadrant splits this line into two axes and is used to capture research more precisely.
The classic three-way classification (OECD, 2015)
Basic research: the pursuit of knowledge without any direct application in view. Examples: elucidating the double-helix structure of DNA; observing gravitational waves (at the moment of discovery, knowledge acquired without regard to its use)
Applied research: the pursuit of knowledge directed at a specific practical goal. Examples: research to make mRNA technology work as a vaccine; drug-discovery research targeting a specific cancer
Development: activity that embodies knowledge in products and processes. Examples: establishing a mass-production process for mRNA vaccines; designing and prototyping a new smartphone
Pasteur's quadrant (Stokes, 1997): four quadrants defined by two axes—"does it seek understanding of fundamental principles?" and "is it conscious of practical use?"
Bohr type: the pure pursuit of understanding. Seeks to understand principles, with no eye to practical use. Examples: the construction of quantum mechanics, epitomized by Bohr's atomic model; cosmology probing the origin of the universe; the search for the Higgs boson—all asking only "how is the world?" regardless of any use
Pasteur type: the simultaneous pursuit of understanding and use. Seeks to understand principles while keeping practical use in view. Example: Pasteur worked out "why" fermentation and putrefaction occur, founding microbiology as a basic discipline, while at the same time producing vaccines and pasteurization as practical results. Bell Labs' solid-state physics research (where elucidating the principles of semiconductors led directly to the invention of the transistor) is of the same type. Today's AI research, which seeks to understand the principles of intelligence while generating enormous industrial applications, is a textbook case of this quadrant
Edison type: the pure pursuit of use. Does not seek to understand principles, but is conscious of practical use. Example: Edison tried thousands of filament materials for the light bulb, yet showed no interest in explaining the physics of incandescence. Much product-improvement research that pushes performance within the bounds of known principles falls here
The fourth quadrant: systematic observation and recording of particular phenomena. Seeks neither to understand principles nor to serve practical use. Example: Stokes himself cited natural-history observation and classification, typified by Peterson's field guides to birds. This work does not directly aim at theory building or practical use, but it accumulates the data that later research builds on
2.4 Three Types of Research Approach: Classification by Research Method
All research is carried out through three modes of inference: deduction, induction and abduction. Inference can be written as a combination of the following three elements (Peirce, 1878).
Rule: a general connection of the form "if A, then B." Bean example: "Any bean taken from this bag is white."
Case: the fact that a particular object stands on the input side (A) of that rule. Bean example: "This bean was taken from this bag." In other words, the reason or origin of its being white.
Result: what actually appears on the output side (B) for that particular object. Bean example: "This bean is white." In other words, the visible fact.
The difference between the three modes of inference lies in which two of these elements serve as the raw material and which remaining one is derived.
Deduction: deriving the Result from Rule + Case
Bean-bag example: all the beans in this bag are white (Rule) + this bean was taken from this bag (Case) → this bean is white (Result)
Examples: mathematically deriving the observed value of "the precession of Mercury's perihelion" (Result) from general relativity (Rule); proving theorems from the axioms of Euclidean geometry (Rule)
Induction: deriving the Rule from Case + Result
Bean-bag example: every bean taken from this bag (Case) has been white (Result) → all the beans in this bag are probably white (Rule)
Example: from Tycho Brahe's vast body of astronomical observations (an accumulation of Cases and Results), Kepler derived the three laws of planetary motion (Rule). Pattern discovery by machine learning is the modern version of this
Abduction: deriving the Case from Rule + Result
Bean-bag example: all the beans in this bag are white (Rule) + there is a white bean on the desk (Result) → perhaps this bean was taken from that bag (a hypothesis about the Case)
Examples: the discrepancy between Uranus's orbit and calculation (Result) was set against the law of universal gravitation (Rule), yielding the Case hypothesis that "an unknown planet is exerting a gravitational pull," which led to the discovery of Neptune. The discovery of penicillin likewise began with a Case hypothesis to explain "bacteria are dying only around the mold (Result)"
In research, the three modes of inference are used in a division of labor. Abduction invents a Case hypothesis; deduction derives testable Results (predictions) from that hypothesis; and induction accumulates observations to establish it as a Rule. Research advances through this cycle.
Science has so far passed through five paradigms (Gray, 2009), and that history is one of turning the three modes of inference, one by one, into tools—that is, mechanizing them.
Paradigm 1 (experimental and observational science): observing and experimenting on nature and discovering laws from experience (a history spanning millennia)
Humans run all three modes of inference by hand (with induction as the mainstay)
Abduction (by hand): in Darwin's evolutionary hypothesis, organisms possess traits adapted to their environment (Rule) + the shape of finches' beaks differs from island to island (Result) → perhaps species have diverged in response to the environment of each island (Case).
Deduction (by hand): in Mendel's experimental design, inheritance is transmitted in particles with dominant and recessive forms (Rule) + this cross is the second generation of a cross between pure lines (Case) → the traits should segregate in a 3:1 ratio (Result).
Induction (by hand): in Galileo's law of falling bodies, each trial of dropping balls of various weights (Case) + in every trial the manner of falling did not depend on weight (Result) → the motion of falling bodies is independent of weight (Rule).
Paradigm 2 (theoretical science): explaining and predicting phenomena through mathematical models and deduction (from the 17th century)
Deduction takes the leading role.
Abduction (by hand): in Newton's universal gravitation, a body changes its direction of motion when a force acts on it (Rule) + apples fall and the Moon keeps circling the Earth (Result) → perhaps the Moon and the apple are both subject to the same pull from the Earth (Case).
Deduction (by hand): in Einstein's prediction, general relativity's "mass curves spacetime" (Rule) + during a solar eclipse, starlight passes right beside the massive Sun (Case) → the apparent positions of the stars should shift by 1.75 arcseconds (Result).
Induction (by hand): each opportunity for testing, such as eclipse observations at various locations and measurements of Mercury's orbit (Case) + every one of them matched the prediction (Result) → general relativity is accepted as an established law (Rule).
Paradigm 3 (computational science): handling complex systems that cannot be solved analytically through simulation (from the late 20th century)
The mechanization of deduction. Computers take over deductions at a scale that cannot be solved analytically.
Computers carry out the deductive step Rule (equations) + Case (initial conditions) → Result (numerical solution) at scales beyond analytical solution
Deduction (mechanized): In climate modeling, the equations describing the physics of the atmosphere and oceans (Rule) plus the current state of the atmosphere as initial conditions (Case) yield a numerical solution for the temperature distribution 100 years from now (Result).
Abduction (manual): When improving a climate model, a model built on known physical laws (Rule) plus computed temperatures that diverge from observations (Result) led to the inference that the model might be missing the cooling effect of aerosols (Case).
Induction (manual): Computations run from past initial conditions (Case) plus the observed historical climate record (Result) led to refining the model toward the parameter values that best fit the observations (Rule).
Fourth paradigm (data-driven science): discovering patterns in large-scale data through statistics and machine learning (2000s onward)
The mechanization of induction. Statistics and machine learning take over the inductive step of extracting a Rule (a pattern) from large numbers of Case–Result pairs.
Induction (mechanized): In GWAS, the genomes of hundreds of thousands of individuals (Case) plus whether each of them has a given disease (Result) yield a pattern linking specific genetic variants to the disease (Rule).
Deduction (mechanized): In prediction with a trained model, the pattern acquired during training (Rule) plus a new patient's genome (Case) yield a prediction of that patient's risk of developing the disease (Result).
Abduction (manual): In elucidating the mechanism, the correlation between variant and disease (Rule) plus observations of gene expression in patient tissue (Result) led to the hypothesis that this variant causes the disease through this particular mechanism (Case).
The fourth paradigm ran into a characteristic problem: mechanical induction churns out correlations (Rules) from data (Cases and Results) in vast quantities, but explanation (the process that connects Case to Result) cannot keep up.
For example, in 2007, data on the genomes of hundreds of thousands of people (Case) and whether each of them was obese (Result) yielded the correlation (Rule) that people carrying a variant in the FTO region are more prone to obesity. But the step from this correlation and the obesity in front of researchers to the hidden cause—the mechanism by which the variant produces obesity—stayed unfilled for eight years (abduction was still done by hand). Only in 2015 did it emerge that the FTO variant alters the expression of IRX3/IRX5 and thereby reduces heat production in fat cells.
Cases and Results form a chain (e.g., Case 1 → Result 1 = Case 2 → Result 2 = Case 3 → Result 3 ...), and unraveling that chain is what constitutes an explanation of the Rule. In the FTO example, the chain runs: Case 1 = variant in the FTO region → Result 1 = increased IRX3/IRX5 expression (= Case 2) → Result 2 = fat precursor cells differentiate into the fat-storing type rather than the heat-producing type (= Case 3) → Result 3 = reduced energy expenditure (= Case 4) → Result 4 = fat accumulation (= Case 5) → Result 5 = obesity. What GWAS found was only a coarse Rule connecting the two ends (Case 1 and Result 5) directly; to "explain" is to fill in this chain of intermediate Cases and Results and to ground each link in a known Rule (enhancers alter gene expression, changes in expression alter cell differentiation, and so on). Because the intermediate Cases and Results (expression levels, cell states and the like) are not observed in the data and only become visible through new experiments, the problem was that the manual work of elucidating the middle could not keep pace with the speed at which the outer correlations were being discovered.
Fifth paradigm (AI-driven science): AI runs the research loop itself, from hypothesis generation through experimental design, execution and interpretation (Microsoft Research, 2022; Leng et al., 2023)
The mechanization of all three modes of inference, abduction included.
Abduction (mechanized): In AI-driven hypothesis generation, the body of known laws in the literature (Rule) plus unexplained experimental data (Result) lead the AI to propose a new mechanistic hypothesis (Case). Example: Google's AI co-scientist (2025) has independently proposed mechanistic hypotheses consistent with experimental results for an open problem in bacterial gene transfer (Gottweis et al., 2025).
Deduction (mechanized): In AlphaFold's structure prediction, the learned sequence-to-structure mapping (Rule) plus the amino acid sequence of a given protein (Case) yield a prediction of that protein's three-dimensional structure (Result).
Induction (mechanized): In training AlphaFold, hundreds of thousands of amino acid sequences (Case) plus experimentally determined three-dimensional structures (Result) yield the mapping between sequence and structure (Rule).
2.5 Classifying Research Approaches (3): By Research Output
Research can also be classified by the type of output it produces—that is, by what kind of knowledge it adds.
New theory: a new explanatory framework for phenomena
Examples: special relativity (Einstein, 1905), the DNA double-helix model (Watson & Crick, 1953), and "The Market for Lemons," which showed how information asymmetry leads to market failure (Akerlof, 1970)
New method: new means of measurement, analysis or construction
Examples: PCR for amplifying DNA fragments (Mullis & Faloona, 1987), CRISPR-Cas9 genome editing (Jinek et al., 2012), and the Transformer, which became the foundation of generative AI (Vaswani et al., 2017)
New evidence: empirical findings that support or refute an existing theory
Examples: the eclipse observations that tested general relativity (Dyson, Eddington & Davidson, 1920), the epidemiological study linking smoking to lung cancer (Doll & Hill, 1950), and the first direct detection of gravitational waves (Abbott et al., 2016)
New artifacts: systems, datasets, benchmarks
Examples: ImageNet, the image dataset that laid the groundwork for the deep learning boom (Deng et al., 2009), and the AlphaFold Protein Structure Database, which released more than 200 million protein structures (Jumper et al., 2021)
Synthesis: surveys, meta-analyses, systematic reviews
Examples: the synthesis of psychotherapy outcome studies that established meta-analysis itself as a method (Smith & Glass, 1977), and the IPCC assessment reports, which integrate climate research from around the world into a shared foundation for policy
Replication: testing whether existing findings can be independently reproduced
Examples: the large-scale project that systematically replicated 100 prominent psychology studies and found that only about 36% reproduced with significant results (Open Science Collaboration, 2015), and the replication of 21 social science studies published in Nature and Science (Camerer et al., 2018)
3. Who Does Research
3.1 How the Actors in Research Have Changed
Over time, the central actors in research have shifted as well.
Through the 19th century: individuals and patrons — Research was largely the pastime of aristocrats, or was carried out under the protection of princes and the Church
Examples (researchers): Galileo enjoyed the patronage of the Medici family and repaid them by naming the moons he discovered orbiting Jupiter the "Medicean Stars." Cavendish, an aristocrat, used his private fortune to discover hydrogen and measure the density of the Earth, and Darwin too was a "gentleman scientist" who funded his own research
Example (patron): Tycho Brahe was granted an island and funding by the King of Denmark, and built what was then the finest observatory in the world
19th century onward: universities — The Humboldtian university turned research into a profession and made universities the center of knowledge production
Examples (institutions): Liebig's laboratory at the University of Giessen invented modern laboratory-based training, putting students to work on experiments, and turned out chemists in large numbers. The Cavendish Laboratory at Cambridge (founded 1874), beginning with its first director Maxwell, went on to produce J.J. Thomson (discovery of the electron) and Rutherford (discovery of the atomic nucleus)
Examples (researchers): Pasteur, Koch and others like them—researchers by profession, affiliated with universities and institutes—became the norm in this era
Early to mid-20th century: corporate central research labs and government
Examples (corporate): Bell Labs (the transistor by Shockley and colleagues, Shannon's information theory—and multiple Nobel Prizes), Xerox PARC (invention of the GUI and Ethernet), DuPont (Carothers's nylon), and the GE Research Laboratory
Examples (government): the Manhattan Project (Los Alamos National Laboratory under Oppenheimer), NASA's Apollo program, and the founding of CERN (1954)—government-led big science expanded
Late 20th century: corporate labs shrink and roles divide — Under pressure for short-term profits, companies retreated from basic research, and a division of labor emerged in which universities handle the fundamentals and startups commercialize the results (in the US, the Bayh-Dole Act of 1980, which allowed universities to patent their inventions, gave this a push)
Examples (symbols of the retreat): the decline of Bell Labs after the breakup of AT&T. Xerox PARC's inventions being commercialized not by Xerox itself but by Apple and others—the classic "one that got away"—is regarded as emblematic of the limits of the central research lab model
Examples (university spinouts): Genentech (1976), founded by UCSF's Boyer and colleagues, became the prototype for the biotech startup, and Stanford students Page and Brin founded Google (1998)
Today: the rise of startups and new research organizations — "New research organizations" is an umbrella term for a new breed of research organization that is neither a university, a government agency nor a conventional company
Examples (research-focused startups): OpenAI and Anthropic (AI), SpaceX (space), Moderna and BioNTech (mRNA). In AI especially, organizations founded only a few years ago are leading frontier research itself
Examples (new types of nonprofit): FROs, time-limited organizations that solve a single scientific bottleneck and then disband, and Arc Institute, which gives researchers long-term, unconditional funding so that they need not apply for grants
With the rise of startups and new research organizations, the following trends have gathered pace in recent years.
Companies and startups are leading research
In AI, roughly 90% of notable AI models now come from industry (Stanford HAI, 2025). Leadership of frontier research has passed from universities to organizations such as OpenAI, Anthropic and Google DeepMind.
By making reusable rockets a reality, SpaceX rewrote, by orders of magnitude, the cost structure of a launch sector that national space agencies had dominated for half a century.
Moderna and BioNTech concentrated their investment on mRNA, a technology that had long remained unproven, and delivered COVID-19 vaccines faster than any vaccine in history
The scale of computing resources, data and capital now required has outgrown what universities can procure. Training a foundation model calls for investment on the order of tens of billions of yen—orders of magnitude beyond the world of conventional government grants
In the major economies, the business sector accounts for roughly 70% of R&D spending, and the R&D budget of a single tech giant exceeds the government research budgets of many countries
Experiments with research organizations that are neither universities nor conventional companies—such as FROs (Focused Research Organizations) and nonprofit institutes like Arc Institute—are also gaining momentum
Open versus closed is increasingly a matter of selective design
The more the center of research shifts to companies and startups, the greater the tension between the ethos of science (the norm of openness, discussed below) and the competitive logic of business (keeping things closed)
Increasingly, the technical details of frontier models are never published as papers at all. Research results that "don't even appear as preprints" are on the rise.
The selective design of what to open and what to keep closed is no longer merely a matter of corporate strategy; it has become an institutional question of how to protect knowledge sharing across the research system as a whole.
3.2 Comparing the Characteristics of Research Actors
We compare today's principal research actors along the following four dimensions.
Funding source: whose money pays for the research, and at what scale
Center of gravity: whether the emphasis lies on basic or applied research
Speed: the agility of decision-making and of the research cycle
Openness: how far results are made public
Using these four dimensions, we compare six types of actors.
Universities
MIT, Stanford University, the University of Oxford, and others
Their strengths are the integration of talent development with research and the long-term, unconstrained pursuit of questions. Their constraint is funding scale: their capacity to procure large-scale compute and facilities lags behind that of companies.
Funding source: competitive government grants, block operating subsidies, and donations. Individual awards are small.
Center of gravity: the heartland of basic research, and the only type of actor that covers every field
Speed: slower (rate-limited by grant cycles and the dual burden of teaching)
Openness: the highest. Publishing papers is itself the mission and the measure of achievement
National Research Institutes and Government Labs
NIH (the U.S. National Institutes of Health), CERN, NASA, Los Alamos National Laboratory, and others
Their strengths are massive infrastructure that universities cannot own (accelerators, supercomputers, large telescopes) and continuous observation sustained over decades. Their constraints are low agility and exposure to political cycles.
Funding source: government budgets. Stable and large-scale
Center of gravity: from basic through applied. They carry national missions (security, public health, standards-setting)
Speed: slow in ordinary times (bureaucracy), but capable of overwhelming concentration of effort on national priorities (e.g., Manhattan Project–style mobilization)
Openness: high (though security-related and sensitive areas are classified)
Corporate Research Laboratories
Bell Labs, Google Research, Microsoft Research, IBM Research, and others
Their strengths are vertical integration from research through to product, and the scale of their data and facilities. Their constraint is short-term profit pressure, which tends to drive them out of basic research (as the history of shrinking central research labs shows).
Funding source: the company's own revenue, tied to the business cycle and management priorities
Center of gravity: primarily applied, plus selective basic research (aimed at maintaining the absorptive capacity to take in external science; discussed later)
Speed: fast (funding decisions are made entirely within the organization)
Openness: selective. What they sell (substitutes) is kept closed; what drives adoption (complements) is opened up
Startups and Emerging Research Companies
OpenAI, Anthropic, SpaceX, Moderna, and others
Their strength is the breakthrough power that comes from single-minded focus. Beyond leading frontier AI research, SpaceX used reusable rockets to break into territory long monopolized by national space agencies, and Moderna turned mRNA, then an unproven technology, into a working vaccine. Their constraints are the risk of running out of money and a short time horizon: when they fail, their accumulated knowledge tends to scatter with them.
Funding source: venture capital and large investors. Enormous amounts of risk capital concentrated on a single bet
Center of gravity: from applied through Pasteur's quadrant. All research resources are thrown behind a single mission
Speed: the fastest. Decisions are made in days, with no deadlines and no grant cycles
Openness: shifts strategically (they tend to be open early on to attract talent, then close up once an advantage is established)
Nonprofit and Private Research Institutes / Foundations
Howard Hughes Medical Institute, the Allen Institute, Arc Institute, and others
Their strength is long-term, high-risk research unconstrained by the grant system (the HHMI-style model of "investing in people, not projects"). Their constraints are the limits of scale and dependence on funders' wishes.
Funding source: endowments and donations. Freer than government, longer-term than companies
Center of gravity: from basic research through high-risk areas. Their raison d'être is taking the risks that government grants will not
Speed: moderate to fast (little bureaucracy)
Openness: high (many make it their mission to provide databases and tools as public goods, e.g., the Allen Brain Atlas)
Independent Researchers and Citizen Scientists
Independent scholars, participants in citizen science projects, and others
Their strength is the freedom to pursue questions unconstrained by institutions, with a substantial track record to show for it (mathematics, comet discoveries by amateur astronomers, contributions to open-source software, and more). Their constraints are funding and equipment, and the lack of the credibility signal that an institutional affiliation provides.
Funding source: personal funds, crowdfunding, small grants
Center of gravity: areas with light capital requirements, such as theory, observation, and data analysis
Speed: the fastest (no organizational coordination whatsoever is required)
Openness: high (published directly via preprints, blogs, and open-source software)
Viewed across these four dimensions, one thing stands out: the nature of the funding—whose money it is—largely determines time horizon, center of gravity, and openness. Government money tends toward the long-term, the basic, and the open; market money toward the short-term, the applied, and the selectively open; and foundation money toward the high-risk territory in between.
3.3 Researcher Evaluation and Careers
This section lays out how researchers are evaluated (the mechanisms that determine hiring, promotion, and the allocation of research funding) and the realities of research careers.
What have researchers been judged on?
Quantity and quality of output: number of papers, citation counts, the h-index (a composite measure of productivity and impact), and the prestige of the venue (impact factor, acceptance at top conferences)
Funding acquisition: a track record of winning competitive research grants. This is the lifeline of running a lab, and is itself counted as an achievement
Contribution to the community: mentoring students, peer review, running academic societies. Important, but hard to capture in quantitative metrics and therefore prone to being undervalued
The problem is that the side effects of metric dependence have become the norm: "publish or perish" pressure, slicing results into the thinnest publishable pieces (salami publishing), and overreliance on impact factors. International soul-searching over this fixation on metrics, such as the DORA declaration (2013), is now being institutionalized as well
The Standard Career Path and Its Structural Problems
The standard path runs: PhD → postdoc (fixed-term researcher) → tenure track (the probationary period before a permanent appointment) → tenure → PI (principal investigator running a lab)
The structural problem is that the number of positions has not kept pace with the growth in PhD holders. Postdocs are getting longer and more precarious worldwide, and tenure has become a narrow gate
For the majority of PhD holders described above, careers outside academia (in industry, government, and startups) are gradually becoming the norm
The Talent Shift to Industry and the Pay Gap
In AI, roughly 70% of PhD graduates now go into industry (20 years ago the figure was around 20%; the ratio has flipped)
Top AI researchers are offered compensation on the order of hundreds of millions of yen a year by startups and big tech, and media reports describe offers worth billions of yen in total contract value (e.g., the poaching war that accompanied Meta's establishment of its superintelligence lab) (CNBC, 2025). That is two orders of magnitude above a university faculty salary.
"Acqui-hires"—acquisitions made for the purpose of obtaining talent, in which securing outstanding researchers or research teams stands in for buying a company—have also become commonplace. Cases of entire research teams moving en masse keep coming.
As for the impact on universities, they are becoming institutions that "train talent but cannot retain it," raising concern about the hollowing out of the faculty ranks that educate the next generation. At the same time, a "revolving door" pattern of mobility is emerging, in which research experience gained in industry flows back into universities
4. The Institutions of Research
4.1 Sharing Research
The history of academic research can be read as a history of inventing the institutions that support the fourth of the four constituent acts of research described earlier: sharing. The institutions for sharing research have traced the following history.
Prehistory: The Age of Secrecy (to the 17th Century)
The structural problem: publishing a discovery brought no benefit, only the loss of having it stolen. The more rational the scholar, the more likely they were to hide their discoveries
Examples: Galileo recorded his astronomical discoveries as anagrams (ciphers made by rearranging letters), and Hooke published his law of springs as the cipher "ceiiinosssttuv" (later revealing the solution: "ut tensio, sic vis"—as the extension, so the force)
Example: in the alchemical tradition, concealing results behind jargon and codes was standard practice
17th Century: The Invention of Journals and Priority Rules
With the spread of printing, the Royal Society of London, founded in 1660, launched Philosophical Transactions in 1665. The same year saw the launch of the Journal des sçavans in Paris, and the academic journal as a format was born
A rule was established whereby priority is recognized not by hiding a discovery but by publishing it (rather than being allowed to monopolize the discovery, the discoverer gets their name attached to it—the honor of being the first to find it). The incentive to hide knowledge was flipped into an incentive to share it
Journals thereby acquired four functions that survive intact to this day: registration (the official record of who discovered what and when, i.e., the guarantee of priority), certification (assurance that the content is sound), dissemination (communication to the community), and archiving (preservation for posterity)
19th Century: Professionalization and Specialization
The Humboldtian university, beginning with the University of Berlin in 1810, championed the "unity of teaching and research" and turned research from a pastime (a hobby or gentlemanly accomplishment) into a profession
As disciplines specialized, specialist societies and journals split off from general learned-society journals, and the PhD and the seminar system became the standard for training researchers
This German model was exported to the United States (Johns Hopkins University and others) and became the prototype of today's research university
20th Century: Public Funding and the Institutionalization of Peer Review
Vannevar Bush's report Science, The Endless Frontier (Bush, 1945) set in motion large-scale government funding of research and the establishment of funding agencies such as the NSF and NIH
Systematic peer review spread across academic journals in general only from the mid-20th century onward: rising submission volumes and specialization pushed individual editors' judgment to its limits, prompting a shift to review divided among external experts
Example: in 1936, when Physical Review sent a paper Einstein had submitted out for external review, he protested that he had not authorized it to be shown to others and withdrew the paper. Even for a leading physicist of the day, peer review was still a novel mechanism.
Late 20th Century to the Present: Digitization Expands Sharing
As digitization advanced, arXiv, launched in 1991, made preprints the standard in physics, and the open access movement of the 2000s spread a publishing model in which readers can read for free
With permanent identification via DOIs and the normalization of publishing data and code, the unit of "sharing" is expanding from the paper to data, code, and models
The greatest asset produced by this history of sharing institutions is science as a domain of cooperation that transcends national borders, supported by two layers: norms and institutions.
The Normative Layer: The Ethos and Values of Science (Merton, 1942)
Communalism: results are the common property of the community
Universalism: work is judged on its content, not on nationality or status
Disinterestedness: truth is pursued independently of personal interest
Organized skepticism: every claim is subjected to systematic scrutiny
Combined with the priority rule seen in this chapter—under which publication counts as achievement—these norms gave real force to knowledge sharing across borders
The Institutional Layer: Mechanisms of International Cooperation
International mutual aid in peer review: researchers in other countries review work for free
The expansion of international co-authorship and large international projects: CERN, ITER, the Human Genome Project
Example: the Human Genome Project's "Bermuda Principles" (1996) required sequenced DNA to be released within 24 hours—a landmark that deliberately designed the pre-competitive knowledge base as a commons for all humanity
A foundational knowledge base is a complement for every stakeholder—companies, governments, and individuals alike—so there is a rational case for making it a commons
4.2 How Research Is Evaluated
The peer-review criteria used by academic journals and international conferences—that is, how research is evaluated—are worded differently across fields and venues, but they share a common core that boils down to the following five points.
Novelty / originality: What does the work add to existing knowledge?
Significance: How much does that new contribution matter to the field or to society?
Soundness / rigor: Are the claims properly supported by the methods and evidence?
Clarity: Is it written so that readers can understand and verify it?
Fit: Does it suit the readership and purpose of the venue?
These five criteria have a known weakness, however. In what has been called the "novelty penalty" (documented in a study analyzing more than 20,000 peer reviews across 49 journals; Teplitskiy et al., 2022), journals say they want highly novel research, yet the more novel a paper is, the harsher the reviewers tend to be. Highly novel work does not fit existing yardsticks and imposes high comprehension and verification costs on reviewers, which pushes evaluations toward risk aversion. There are many well-known stories of research that later won a Nobel Prize being rejected by top-tier journals.
4.3 Types of Publication Channels
The fourth act of research, sharing, can take place through several channels. We compare them along the following six dimensions.
Review (assurance of accuracy): Who checks the content, and how rigorously?
Speed: How long does it take from a completed result to public release?
Cost: The financial burden and effort required of the authors
Reach and access: Who does it reach, and how openly?
Academic credit: How it counts in researchers' careers, hiring, and promotion
Permanence and citability: Is it preserved permanently as a definitive version, and can it be cited?
Using these six dimensions, we compare six major channels.
Peer-reviewed academic journals
e.g., Nature, Science, PNAS, American Economic Review
Review: Systematic peer review by external experts—the strongest quality assurance of the six channels
Speed: Slow (months to years from submission to publication)
Cost: Subscription journals charge nothing to submit (readers bear the cost). Open-access journals charge authors an article processing charge (APC) that can run to several hundred thousand yen
Reach: Mainly through subscriptions and libraries. Open-access journals can be read by anyone
Academic credit: The most important channel in medicine, the natural sciences, and much of the social sciences
Permanence: Preserved permanently with a DOI as the "version of record"
International conferences (peer-reviewed proceedings)
e.g., NeurIPS, ICML, ACL
Review: Peer-reviewed. Selectivity itself (acceptance rates of roughly 20–30%) serves as a quality signal. However, the review of each individual paper is lighter than at journals, and this is also the channel where the reviewer overload described at the start of this article is most acute
Speed: Fast (months from submission to presentation; deadline-driven)
Cost: Authors bear substantial registration and travel costs
Reach: Beyond the published proceedings, presentations and discussions reach peers directly
Academic credit: In computer science, equal to or greater than journals. In most other fields, supplementary—the core of a research record remains journal articles
Permanence: Preserved as proceedings with a DOI
Preprints
e.g., arXiv, bioRxiv, SSRN
Review: None (formal checks only). Assurance of accuracy is left entirely to readers' ability to verify
Speed: Fastest (same day to a few days)
Cost: Free
Reach: Free for anyone to read—the most open of the six channels
Academic credit: Establishes priority, but counts only provisionally as a research achievement (standard in physics and AI; varies widely across other fields). Example: starting with the Transformer paper (Vaswani et al., 2017), AI research is effectively preprint-driven
Permanence: Permanently archived with a DOI, but not a peer-reviewed "version of record"
Books and monographs
e.g., Pasteur's Quadrant (Stokes, 1997), Piketty's Capital in the Twenty-First Century
Review: By the publisher's editors and external reviewers (looser than peer review, with varied standards)
Speed: Slowest (years to write and publish)
Cost: The greatest writing effort of the six channels. Academic books may also require publication subsidies
Reach: Virtually the only channel that reaches general readers beyond the research community
Academic credit: Most important in the humanities and social sciences; supplementary in the sciences
Permanence: Read for decades; the longest-lasting influence of any channel
University archives, institutional repositories, and dissertations
e.g., the Stanford University and Harvard University research archives
Review: Dissertations are examined by a committee. Technical reports receive only internal checks
Speed: Fast (only the university's own procedures are required)
Cost: Free (operated by the university)
Reach: Accessible to anyone, but hard to search for and discover
Academic credit: Rarely counts as an achievement in itself (though a dissertation is a prerequisite for becoming a researcher). Example: PageRank, the foundation of Google Search, was released as a Stanford University technical report and became a historic contribution (Page et al., 1999)
Permanence: Permanently preserved by the university (the "green" open-access route)
Releasing data, code, and models
e.g., GitHub, Zenodo, Hugging Face
Review: Essentially none (though peer-reviewed "data papers" have emerged in recent years)
Speed: Same day
Cost: Free to low-cost
Reach: Grows through reuse. Example: ImageNet (Deng et al., 2009) and the AlphaFold Protein Structure Database (Jumper et al., 2021) have generated more citations and use than most papers
Academic credit: Rising in status as an independent contribution, but institutional recognition is still a work in progress
Permanence: DOIs are becoming common, but the risks of broken links and abandoned maintenance remain
Viewed across these six dimensions, channel choice is fundamentally a matter of trade-offs.
Review rigor and speed are inversely related: the stronger the quality assurance, the slower the channel (journals, books); the faster the channel, the more verification is left to readers (preprints, data releases)
The weight given for academic credit is set by field conventions: the same channel can carry the opposite weight in computer science (conferences), the humanities and social sciences (books), and medicine and the natural sciences (journals)
In practice, channels are combined: rather than relying on a single channel, the emerging standard is to share quickly via preprint → certify through journals and conferences → ensure verifiability by releasing data and code
Fields also differ in which publication channels they prioritize. This can be explained by the fit between the nature of a field's knowledge and the characteristics of each channel (review rigor, speed, and capacity). Four factors determine which channels a field favors.
① The size of the unit of knowledge
Medicine and the natural sciences: The unit of contribution is "one experiment, one discovery," which fits neatly into a single paper → the journal is the natural vessel
Humanities and social sciences: The unit of contribution is a system of interpretation or a historical narrative that cannot fit within the length of a paper. It becomes persuasive only as a structure spanning hundreds of pages → the book is the natural vessel
Computer science: Progress comes in small, rapid increments, suiting a style of frequent short papers → the conference is the natural vessel
② The pace of change in the field
Computer science: Technology becomes obsolete quickly, and waiting through a journal's one- to two-year review cycle would leave results stale → deadline-driven conferences that publish within months match the demand for speed
Humanities and social sciences: Findings age slowly, and being read for decades is what confers value → the years it takes to publish a book are not a drawback
③ The social cost of error
Medicine: Wrong findings directly affect treatment and human lives → the journals with the most rigorous review (systematic peer review, statistical review, conflict-of-interest disclosure) remain the standard for academic credit. The demand for rapid communication is met through a division of labor with letters journals such as Nature and Physical Review Letters
④ Historical path dependence
Computer science: The field grew explosively in the 1970s and 1980s. This young discipline had no established hierarchy of prestigious journals yet, so ACM and IEEE conferences were the first to build a credit infrastructure of "rigorous peer review + selection by acceptance rate + published proceedings." Once hiring and promotion decisions began to turn on "how many papers at top conferences," researchers sent their best work to conferences, which further raised conference prestige—a self-reinforcing cycle
Medicine and the natural sciences: 350 years of journal infrastructure (along with evaluation devices such as the impact factor) continue to hold sway
5. Putting Research to Use
5.1 Research Develops in Response to Social Demand
The development of research cannot be explained by the supply-side force of intellectual curiosity alone. Economics offers the classic distinction between "technology-push" and "demand-pull" (Mowery & Rosenberg, 1979), and a look back at history shows that the areas tied to major demand—especially war and social challenges—attracted concentrated funding, talent, and publication volume, and advanced the fastest.
Demand from war and national security
In World War II, the Manhattan Project poured unprecedented resources and talent into physics, radar development into electronics, and codebreaking (the work of Turing and others) into the birth of computer science
During the Cold War, the Sputnik shock led to the founding of NASA and ARPA (now DARPA); ARPANET evolved into the internet, and military positioning systems into GPS
Regions and technology fields that received more wartime R&D investment from the OSRD in World War II saw inventive activity continue to increase for decades after the war (Gross & Sampat, 2023). Government defense R&D has also been shown not to crowd out private R&D investment but to induce it ("crowd in"), raising industrial productivity (Moretti, Steinwender & Van Reenen, 2025)
The Bush report (Bush, 1945) was itself the institutionalization of the wartime research success story; the postwar system of government-funded research was a peacetime transplant of "a mechanism that war had built"
Demand from social challenges
During COVID-19, research resources converged on infectious disease all at once, compressing vaccine development that normally takes ten years into under a year. Related papers also exploded in number, driving a rapid expansion of biomedical preprints (Fraser et al., 2021)
Major social challenges—the Nixon administration's "War on Cancer" (1971), AIDS research, climate change research (the vast body of papers synthesized by the IPCC)—draw funding, talent, and publication volume to their domains. It has also been shown empirically that NIH funding allocation by disease tracks, to some degree, the burden of disease (its social weight) (Gillum et al., 2011)
The relationship between demand and research volume
Where money flows, the number of researchers and papers grows. The direction of research is set not only by pure intellectual interest but by society's priorities (funding allocation). Indeed, NIH funding-call data show that researchers shift their research topics in response to grants earmarked for specific purposes (Myers, 2020). During COVID-19 as well, a substantial share of the scientific community's total output pivoted to infectious disease research within a short period (Hill et al., 2025)
Demand-driven research accelerates results, but it also produces distortions: excessive concentration in fashionable areas (a homogenization of research) and underinvestment in basic fields with no conspicuous demand. An analysis of millions of biomedical papers found that researchers overwhelmingly cluster around already-established, popular topics, and that a conservative strategy of avoiding risky new areas is the dominant one (Foster, Rzhetsky & Evans, 2015).
A "pivot penalty" has also been documented: research that pivots to a topic far from the researcher's own specialty tends to have less impact, meaning that abrupt shifts toward demand come at a cost (Hill et al., 2025).
5.2 Research as a Public Good
Research is sustained by public spending in each country, and the knowledge it produces is used as an international public good. Why is it that research is supported by public investment? Two major reasons stand out.
The public-good nature of knowledge (Nelson, 1959; Arrow, 1962)
Non-rivalry: using knowledge does not deplete it
Non-excludability: knowledge is hard to fence off, so its benefits leak out to society as a whole (spillovers)
Consequence: left to the market alone, knowledge is produced in quantities well below what society needs → this is the economic rationale for supporting it through government funding, universities, and the academic community
The return on investment in research (Jones & Summers, 2020)
Each dollar invested in innovation is estimated to yield, on average, more than five dollars in social benefit—and even under conservative assumptions, the returns clearly exceed the amount invested
On average, public investment in research ranks among the best "bargains" in human history
5.3 How Firms Use Research
There are four main pathways through which research reaches the productive activities of firms.
① Direct use of knowledge (paper → patent → product)
A study analyzing citation chains across millions of patents and papers (Ahmadpoor & Jones, 2017) found that most patents can be traced back to a scientific paper within a few citation steps, and that a substantial share of papers, in turn, connect to future inventions
Inventions grounded in scientific papers command higher market value (Krieger et al., 2024)
② Movement of people
When PhD holders move into industry, knowledge is transferred along with the tacit know-how that never appears in papers—experimental tricks, a record of what failed, an instinct for promising questions
Evidence: where the US biotechnology industry was born was determined almost entirely by the location of star scientists, not by capital or markets. This is the classic empirical demonstration that people are the primary channel of knowledge transfer (Zucker, Darby & Brewer, 1998)
The process by which deep learning went from a peripheral topic in university labs to a core industrial technology is, in itself, a history of researchers changing jobs
③ Transfer of artifacts and infrastructure
University spin-off startups, technology transfer systems, and open-source software. The demand-side economic value of open-source software has been estimated at roughly $8.8 trillion (Hoffmann, Nagle & Zhou, 2024)
Public databases (GenBank, the AlphaFold Protein Structure Database, and others). It has been empirically shown that depositing biological materials in public repositories, where anyone can use them, substantially increases the volume of follow-on research that builds on those findings (Furman & Stern, 2011)
④ Absorptive capacity (Cohen & Levinthal, 1990)
Half the reason firms conduct their own research is to maintain the ability to understand and take in science from outside
Firms that write their own papers and attend conferences can draw on research results from around the world faster and more deeply. This concept explains the raison d'être of corporate research labs
5.4 How Governments Use Research
For governments, which coordinate the economy as a whole, the accumulated body of research is the very foundation of policy, regulation, and public services. Their use of it falls broadly into two categories.
① The basis for policy formation (EBPM)
Evidence-Based Policy Making: the movement toward designing and evaluating policy on the basis of research findings and causal inference rather than intuition or precedent (example: testing the effectiveness of anti-poverty programs through randomized controlled trials (RCTs), a case in which a research method from development economics became the standard for policy evaluation as-is (Banerjee & Duflo, 2011))
In an experiment covering 2,150 Brazilian municipalities, mayors and other policymakers who were shown research evidence actually updated their beliefs and changed their policies—and even expressed a willingness to pay for research results (Hjort et al., 2021) (examples: economics in monetary policy and tax design; epidemiological models in pandemic response)
② The foundation for regulation, standards, and public services
Drug approval (statistical standards for clinical trials), environmental regulation (climate science—the IPCC, for example, is an international mechanism that synthesizes research from many countries to build a shared basis for policy), seismic building codes (earthquake engineering), and, among public services and core state functions, public health, weather forecasting, disaster preparedness, and space and defense—none of these could exist without an accumulated body of research
Dedicated institutions for scientific advice—science advisors, advisory councils, national academies—have been developed to connect research knowledge with public administration. Yet science produces "provisional knowledge with uncertainty attached," while politics and administration demand "definitive judgments within a deadline." This gap in time horizons and in how uncertainty is handled is the perennial sticking point of scientific advice (Gluckman, 2014)
5.5 How Households Use Research
Households, the ultimate consumers, also draw on the accumulated body of research as the basis for their everyday decisions.
Healthcare: the quality of treatment an individual receives is determined through clinical practice guidelines—distillations of systematic reviews of the research literature. Underpinning this is the framework of evidence-based medicine (EBM), in which care is based on the best available research findings rather than on the experience or authority of individual physicians (Sackett et al., 1996). Informed consent—the process in which healthcare providers fully explain a patient's condition and treatment options and the patient agrees to treatment—likewise presupposes that research findings are shared
Education and human capital: textbooks and curricula are research findings that have filtered down into standard knowledge over decades. What households actually buy when they invest in education is the accumulated body of research itself
Individuals as participants: with the spread of open access and citizen science, individuals are expanding their role from recipients of research to participants in it. Example: in Galaxy Zoo, the classification of roughly 900,000 galaxies by more than 100,000 volunteer citizens culminated in peer-reviewed papers, and citizen-science discoveries have continued since (Lintott et al., 2008)
Challenges: research findings are easily simplified or distorted on their way to households, and they compete with misinformation. Indeed, large-scale data show that false news spreads faster, farther, and deeper on social media than true news does (Vosoughi, Roy & Aral, 2018)
In summary, the use of research can be organized around the three economic actors—firms (production), governments (coordination), and households (consumption)—as follows.
Firms (production): convert research into new products, services, and productivity through four pathways—knowledge, people, artifacts, and absorptive capacity
Governments (coordination): use research as the basis for policy, regulation, and public services, and maintain the knowledge infrastructure and international cooperation that markets will not supply
Households (consumption): benefit from research as the foundation for decisions about healthcare, education, and consumption, and are beginning to take part on the production side as well through citizen science
6. Research in the Age of AI
So far, we have examined research from four angles: its content (what the activity consists of), its institutions (how it is shared and evaluated), its actors (who does research and how they make a living), and its use (why society supports it and how it is used). We now take these four angles one at a time and ask what changes in the age of AI.
6.1 The Object of Research in the Age of AI
In the age of AI, the human role is likely to center on "questions and verification."
The center of gravity shifts among the four activities
The activity AI will replace most thoroughly is "② generating candidate answers." In literature review, enumerating hypotheses, implementing experimental code, and drafting manuscripts, AI is already far outpacing humans
The activities AI can least readily replace are "① posing questions" and "③ verifying." Picking out, from an endless stream of generated hypotheses and papers, the questions worth answering and the claims that actually hold up demands deep domain knowledge and judgment
In a blind comparison in which more than 100 NLP researchers and an LLM were each asked to generate research ideas, the LLM's ideas were rated more "novel" than those of the human experts, but less feasible—and when actually executed, their ratings fell. This shows that "generating candidates" is already approaching and surpassing human level, while the value remains in "judging which ideas are promising, and executing and verifying them" (Si, Yang & Hashimoto, 2024)
The center of gravity of a researcher's work shifts from generation to "framing questions" and "evaluation and judgment." This is not wishful thinking but a structural change forced by the relocation of the bottleneck
The fifth paradigm swallows its predecessors
The fifth paradigm (AI-driven science) does not replace the first through fourth paradigms; it contains them as components. Internally, the AI research loop consults theory, runs simulations, learns from data, and ultimately verifies through physical experiment
"Coscientist," an LLM-driven system, autonomously carried out everything from literature search and experimental design to controlling robotic lab equipment, and actually performed chemical syntheses. It was an early demonstration of an AI research loop encompassing everything from planning to physical verification (Boiko et al., 2023)
The value of humans who understand the methodological conventions of each paradigm—the assumptions behind statistical tests, the validation of simulations, experimental design—rises rather than falls. To evaluate AI's output, one must be able to evaluate the methods AI used
The values of the contribution types swap places
What loses value: superficial novelty consisting of an unusual combination of existing elements. This is exactly what AI does best, and the moment it can be mass-produced, its value collapses (the "mass production of novel but low-feasibility ideas" shown by Si, Yang & Hashimoto (2024), cited above, is an early sign)
What gains value: deep novelty—"does it move the field's assumptions or the very way questions are posed?"—and verification-oriented contribution types such as replication and reproduction studies and synthesis (the work of adjudicating what counts as established knowledge)
6.2 The Actors of Research in the Age of AI
Firms and startups will become even more dominant as the leading actors
Scale in computing resources, data, and capital becomes the entry ticket to frontier research, so leadership by well-funded firms and startups will continue
At the same time, the gap in absorptive capacity will widen. Anyone can get an AI that reads papers, but the ability to judge what is worth reading and translate it into one's own organizational context accumulates only in organizations that take part in research themselves
At the opposite pole, AI tools extend the individual's capacity for investigation, analysis, and implementation, and individual researchers and citizen scientists stand to gain the most. The actors of research will spread toward both extremes: giant organizations and individuals
In researcher evaluation and careers, "quantity" breaks down as a signal
The more feasible AI-driven mass production of papers becomes, the less paper counts and citation counts function as signals of ability, and researcher evaluation shifts from "quantity" to "verified contributions" and "the quality of the questions asked"
Evaluative labor—peer review, replication checks, data release—will come to be institutionalized as recognized scholarly output
Multi-track careers moving between universities, firms, startups, and nonprofit research institutes will become the norm, and the very dichotomy of "academia or industry" will fade
Accordingly, the metrics used to evaluate researchers can be expected to shift as follows.
Conventional metrics (proxies for quantity)
Publication count: a proxy for productivity. The length of a publication list translates directly into evaluation.
Citation counts and h-index: proxies for impact. Quantity and influence compressed into a single number.
Journal impact factor: paper quality assessed by proxy through "where it was published."
Grant funding secured: the track record of winning research funding is itself counted as an achievement.
Acceptance at highly selective conferences and journals: having passed the filter is itself the signal.
Metrics for the AI era (direct evaluation of substance and behavior)
Narrative descriptions of contributions: narrative CVs and caps on publication lists (the Royal Society's "Résumé for Researchers"; the DFG's ten-publication limit). Researchers are asked to explain their representative contributions in prose rather than by count.
Verifiability and openness: rates of data and code sharing, preregistration, open science badges (Kidwell et al., 2016).
Making roles visible: recording "who did what" through CRediT makes explicit how each individual contributed to the research.
Credit for evaluation labor: peer review, replication, and data management are currently "invisible unpaid labor." But as AI drives an explosion in submissions while the supply of reviewers stays flat, there are moves to record reviewing work (e.g., Web of Science Reviewer Recognition) and to count it alongside publications in hiring, promotion, and grant decisions. NeurIPS now requires submitting authors to take on reviewing duties.
Replication record and self-correction: a documented history of independent parties successfully replicating your claims, and of promptly correcting errors when they are found, is a hard-to-fake behavioral record that predicts the long-term reliability of a person's claims.
6.3 The Institutions of Research in the AI Era
The five standard criteria of peer review will not disappear, but their weighting and operation will shift in five directions.
① Verifiability and reproducibility: from "bonus points" to "mandatory requirement"
Fluent prose and plausible-sounding descriptions of results are no longer evidence of quality. Submitting and publishing code, data, and experimental logs is becoming a de facto requirement for submission across disciplines.
Precedent: since 2019, NeurIPS has had a code submission policy and a reproducibility checklist, which substantially raised code submission rates — an empirical report showing that "making reproducibility a requirement" actually works as an institutional mechanism (Pineau et al., 2021).
② Quality of the question: greater weight at the desk-reject stage
With submissions doubling and reviewer numbers not growing, stronger screening before papers go out to review is unavoidable. Whether a paper can answer "Why this question? Why now? Why this venue?" in its opening will decide acceptance or rejection.
③ Rethinking the "depth" of novelty
Evaluation is turning toward deep novelty, yet the "novelty penalty" (the deeper the novelty, the worse a paper fares in review) shows that the peer-review mechanism itself is built in a way that filters out deep novelty. What must change, therefore, is the structure of the institution rather than its day-to-day operation. Registered Reports, for example, in which only the "question and method" are reviewed and the acceptance decision is fixed before results are seen, offer a mechanism for judging work on the quality of its question rather than on flashy results. They have already spread to more than 300 journals (Chambers & Tzavella, 2022).
④ Provenance and accountability
Beyond disclosing AI use, authors will be asked whether they can explain and defend their own papers. As NeurIPS's new measures show (such as halting review for authors who fail to fulfill their reviewing obligations), the right to submit is becoming bundled with the duty to contribute to evaluation.
⑤ Growing reliance on signals of trust
Reliance on indicators beyond the content itself — track record, affiliation, preregistration, replication history — will grow. While this brings efficiency, designing the trade-off against the risk of disadvantaging newcomers becomes a challenge.
Experiments have shown that for the identical paper, listing a Nobel laureate as author dramatically raises reviewers' recommendations to accept compared with listing an unknown newcomer (Huber et al., 2022). Greater reliance on trust signals could amplify this status bias.
The structure of publication channels will also change.
The "fast sharing" layer (preprints, open data) will swell further as generation costs fall, and a two-tier structure will come into sharp relief alongside the "slow but trustworthy certification" layer (peer-reviewed journals and conferences).
As AI raises the pace of change across every field at once, the field-specific equilibrium among channels — journals, conferences, books — will itself begin to shift.
In a phrase: the center of gravity of peer review shifts from "reviewing content" to "reviewing trust." Is the work shared in a verifiable form, is its provenance transparent, and can the author be held accountable for it? This amounts to an institutional reinforcement of the fourth act of research, "④ Share and make verifiable."
6.4 The Application of Research in the AI Era
Direct access to knowledge — and the simultaneous spread of misinformation
AI-powered search, summarization, and translation of papers dramatically lowers the cost of delivering research findings directly to government EBPM (evidence-based policymaking) and to households' medical and everyday decisions — shortening the "last mile" that once required expert intermediaries. Example: in a blinded comparison, AI chatbot responses to patients' medical questions were rated higher than physicians' responses on both quality and empathy (Ayers et al., 2023). This is early empirical evidence that the cost of mediating expert knowledge is actually beginning to fall.
At the same time, because AI also lowers the cost of generating "plausible but false information," the mechanisms that distinguish verified knowledge (peer review, clinical guidelines, institutions of scientific advice) become more important, not less. Example: experiments have shown that disinformation tweets generated by GPT-3 are harder to detect than disinformation written by humans (Spitale, Biller-Andorno & Germani, 2023).
Renegotiating international cooperation and security
The demands of research security, together with the widening first-mover advantages that AI confers, are intensifying the pressure toward closure.
The OECD has made the reconfiguration of scientific cooperation under geopolitical pressure a central policy theme (OECD, 2025), and in the field of AI it has been empirically shown that US–China rivalry is actually reshaping international co-authorship networks built up over 25 years (Zhang et al., 2026).
Will the time lag shrink?
How far AI can compress the decades-long lag between basic research and societal impact (the long chain from papers to patents measured by Li, Azoulay & Sampat (2017) and Ahmadpoor & Jones (2017)) remains an unresolved empirical question.
What is certain is that the lag stems not only from "generation" but also from "verification and social acceptance." If only generation speeds up and verification fails to keep pace, the lag will not shrink — and a new risk emerges: the transfer of erroneous knowledge into industry, policy, and daily life.
Across all four perspectives, the same single pattern emerges: generation becomes cheap everywhere, and value concentrates in "questions, verification, and trust."
7. Conclusion
Having surveyed research systematically in this way, we have laid out how it will need to transform in the AI era. In light of these changes, researchers and research institutions are likely to be called upon to respond rapidly.
References
Abbott, B. P., et al. (LIGO Scientific Collaboration and Virgo Collaboration) (2016). Observation of Gravitational Waves from a Binary Black Hole Merger. Physical Review Letters, 116, 061102. — The first direct observation of gravitational waves, a century after Einstein's prediction; a representative example of a "new evidence" type of contribution.
Ahmadpoor, M., & Jones, B. F. (2017). The dual frontier: Patented inventions and prior scientific advance. Science, 357(6351), 583–587. — A large-scale analysis of citation chains between patents and scientific papers, showing that most patents reach a scientific paper within a few steps.
Akerlof, G. A. (1970). The Market for "Lemons": Quality Uncertainty and the Market Mechanism. Quarterly Journal of Economics, 84(3), 488–500. — A representative example of a "new theory" type of contribution, showing that information asymmetry produces market failure. The work was recognized with the Nobel Prize in Economics.
Arrow, K. J. (1962). Economic Welfare and the Allocation of Resources for Invention. In The Rate and Direction of Inventive Activity. Princeton University Press. — A classic of economics showing that, because knowledge is a public good, research will be undersupplied if left to the market.
Ayers, J. W., et al. (2023). Comparing Physician and Artificial Intelligence Chatbot Responses to Patient Questions Posted to a Public Social Media Forum. JAMA Internal Medicine, 183(6), 589–596. — A blinded comparative study showing that AI chatbot responses to patients' medical questions were rated higher than physicians' responses on both quality and empathy. Evidence that the cost of direct access to expert knowledge is falling.
Banerjee, A. V., & Duflo, E. (2011). Poor Economics: A Radical Rethinking of the Way to Fight Global Poverty. PublicAffairs. — The culmination of a body of work that brought randomized controlled trials (RCTs) to the evaluation of anti-poverty interventions, becoming a representative example of evidence-based policymaking.
Boiko, D. A., MacKnight, R., Kline, B., & Gomes, G. (2023). Autonomous chemical research with large language models. Nature, 624, 570–578. — An early demonstration of an AI research loop, in which the LLM-driven system "Coscientist" autonomously carried out everything from literature search to experimental design, control of robotic equipment, and execution of chemical synthesis.
Bush, V. (1945). Science, The Endless Frontier. U.S. Government Printing Office. — The report that served as the blueprint for the system of government support for basic research after World War II.
Camerer, C. F., et al. (2018). Evaluating the replicability of social science experiments in Nature and Science between 2010 and 2015. Nature Human Behaviour, 2, 637–644. — A representative example of a "replication" type of contribution, systematically replicating 21 social science experiments published in top journals.
Chambers, C. D., & Tzavella, L. (2022). The past, present and future of Registered Reports. Nature Human Behaviour, 6, 29–42.: A review of the Registered Reports format, in which the research question and methods are peer-reviewed and the acceptance decision made before any results are seen, and of its adoption by more than 300 journals. A concrete example of designing an institution that rewards the quality of the question rather than the flashiness of the result.
CNBC (2025). AI talent war: Tech giants pay talent millions of dollars.: A news report on the competition to recruit AI researchers, documenting compensation offers ranging from millions to hundreds of millions of dollars.
Cohen, W. M., & Levinthal, D. A. (1990). Absorptive Capacity: A New Perspective on Learning and Innovation. Administrative Science Quarterly, 35(1), 128–152.: The classic paper that introduced the concept of "absorptive capacity"—the ability to recognize, assimilate, and apply external knowledge—and thereby explained why firms conduct research of their own.
Deng, J., et al. (2009). ImageNet: A large-scale hierarchical image database. IEEE CVPR 2009.: Built a dataset of more than 14 million images that became the foundation of the deep learning boom. A leading example of a "new artifact" type of contribution.
DFG (2022). Package of Measures to Support a Shift in the Culture of Research Assessment. Deutsche Forschungsgemeinschaft.: A package of research assessment reforms from the German Research Foundation. An early example of institutionalizing "quality over quantity," including a cap of ten on the number of publications that may be listed in a grant application.
Doll, R., & Hill, A. B. (1950). Smoking and Carcinoma of the Lung. British Medical Journal, 2(4682), 739–748.: The epidemiological classic that established the link between smoking and lung cancer. A leading example of a "new evidence" type of contribution based on observational research.
DORA (2013). San Francisco Declaration on Research Assessment.: An international declaration calling for the journal impact factor not to be used in evaluating individual researchers or papers. Signed by thousands of institutions and individuals, it became the starting point for the assessment reform movement.
Dyson, F. W., Eddington, A. S., & Davidson, C. (1920). A Determination of the Deflection of Light by the Sun's Gravitational Field. Philosophical Transactions of the Royal Society A, 220, 291–333.: Verified a prediction of general relativity through observations of a solar eclipse. A leading example of a "new evidence" type of contribution.
Einstein, A. (1905). Zur Elektrodynamik bewegter Körper. Annalen der Physik, 17, 891–921.: The paper that presented special relativity. A leading example of a "new theory" type of contribution.
Foster, J. G., Rzhetsky, A., & Evans, J. A. (2015). Tradition and Innovation in Scientists' Research Strategies. American Sociological Review, 80(5), 875–908.: An empirical study of millions of biomedical papers showing that conservative strategies—clustering around established, fashionable topics—dominate researchers' choices. Evidence of the homogenization of research.
Fraser, N., et al. (2021). The evolving role of preprints in the dissemination of COVID-19 research and their impact on the science communication landscape. PLOS Biology, 19(4), e3000959.: An empirical study quantifying the rapid expansion of biomedical preprints triggered by COVID-19.
Furman, J. L., & Stern, S. (2011). Climbing atop the Shoulders of Giants: The Impact of Institutions on Cumulative Research. American Economic Review, 101(5), 1933–1963.: An empirical study showing that depositing biological materials in public repositories (biological resource centers) substantially increases their use in subsequent research. Evidence that open infrastructure accelerates the accumulation of knowledge.
Gillum, L. A., et al. (2011). NIH Disease Funding Levels and Burden of Disease. PLOS ONE, 6(2), e16837.: An empirical study showing that the NIH's allocation of research funding across diseases tracks, to a degree, the burden of disease (its weight on society). Evidence that societal problems attract funding.
Gluckman, P. (2014). Policy: The art of science advice to government. Nature, 507, 163–165.: An essay by New Zealand's Chief Science Advisor to the Prime Minister framing the gap between scientific uncertainty and governmental decision-making as the central challenge of science advice.
Gottweis, J., et al. (2025). Towards an AI co-scientist. arXiv:2502.18864.: Reports that Google's multi-agent AI system independently proposed a mechanistic hypothesis for an unsolved problem in bacterial gene transfer that was consistent with experimental results. An early demonstration of mechanizing abduction (hypothesis generation).
Gray, J. (2009). Jim Gray on eScience: A Transformed Scientific Method. In The Fourth Paradigm: Data-Intensive Scientific Discovery. Microsoft Research.: The transcript of the talk that formalized the "fourth paradigm" of data-intensive science, following experiment, theory, and computation. The original source of the first-through-fourth-paradigm framework.
Gross, D. P., & Sampat, B. N. (2023). America, Jump-Started: World War II R&D and the Takeoff of the US Innovation System. American Economic Review, 113(12), 3323–3356.: An empirical study showing that regions and technology fields that received wartime R&D investment (via the OSRD) during World War II saw long-lasting increases in postwar inventive activity. Evidence that wartime demand advances research.
Hill, R., Yin, Y., Stein, C., Wang, D., & Jones, B. F. (2025). The pivot penalty in research. Nature.: A large-scale empirical study documenting the "pivot penalty": when researchers shift to topics outside their specialty, the impact of their work declines. Includes an analysis of the massive research pivots that occurred during COVID-19.
Hjort, J., Moreira, D., Rao, G., & Santini, J. F. (2021). How Research Affects Policy: Experimental Evidence from 2,150 Brazilian Municipalities. American Economic Review, 111(5), 1442–1480.: A large-scale experiment showing that policymakers (mayors and other officials) update their beliefs and actually change policy when presented with research evidence. Direct evidence that evidence-based policymaking works.
Hoffmann, M., Nagle, F., & Zhou, Y. (2024). The Value of Open Source Software. Harvard Business School Working Paper 24-038.: Estimates the demand-side economic value of open source software at roughly $8.8 trillion. Evidence of the economic impact of research-derived artifacts (software).
Huber, J., et al. (2022). Nobel and novice: Author prominence affects peer review. PNAS, 119(41), e2205779119.: A large-scale experiment demonstrating "status bias": reviewers' acceptance recommendations for the very same paper change dramatically depending on whether the named author is a prominent researcher or a newcomer.
Jinek, M., et al. (2012). A Programmable Dual-RNA–Guided DNA Endonuclease in Adaptive Bacterial Immunity. Science, 337(6096), 816–821.: Demonstrated the principle of genome editing with CRISPR-Cas9. A leading example of a "new method" type of contribution, and the work recognized by the Nobel Prize in Chemistry.
Jones, B. F., & Summers, L. H. (2020). A Calculation of the Social Returns to Innovation. NBER Working Paper 27863.: Estimates that each dollar invested in innovation yields, on average, more than five dollars in social benefits, demonstrating the magnitude of the social return on research investment.
Jumper, J., et al. (2021). Highly accurate protein structure prediction with AlphaFold. Nature, 596, 583–589.: Raised protein structure prediction to a practically useful level and, by releasing the results as a public database, delivered a "new method" and a "new artifact" at the same time. Recognized by the Nobel Prize in Chemistry.
Kidwell, M. C., et al. (2016). Badges to Acknowledge Open Practices: A Simple, Low-Cost, Effective Method for Increasing Transparency. PLOS Biology, 14(5), e1002456.: An empirical study showing that after the journal Psychological Science introduced open-data badges, the rate of data sharing rose from roughly 3% to roughly 40%. Evidence that behavior changes when openness is made visible as an achievement.
Krieger, J. L., et al. (2024). Standing on the shoulders of science. Strategic Management Journal, 45(6). Demonstrates empirically that inventions grounded in scientific papers command higher market value.
Leng, C., et al. (2023). Fifth Paradigm in Science: A Case Study of an Intelligence-Driven Material Design. Engineering, 24, 126–137. Uses materials design as a case study to give concrete shape to the "fifth paradigm" of AI-driven science.
Li, D., Azoulay, P., & Sampat, B. N. (2017). The applied value of public investments in biomedical research. Science, 356(6333), 78–81. An empirical study drawing on 27 years of data to show that roughly 30% of NIH grants produce papers later cited in private-sector patents. Evidence for the pathways—and the time lags—by which public research investment reaches industry.
Lintott, C. J., et al. (2008). Galaxy Zoo: morphologies derived from visual inspection of galaxies from the Sloan Digital Sky Survey. Monthly Notices of the Royal Astronomical Society, 389(3), 1179–1189. A landmark achievement in citizen science: classifications of roughly 900,000 galaxies by more than 100,000 volunteers, distilled into a peer-reviewed paper.
Merton, R. K. (1942). The Normative Structure of Science. The classic formulation of the "ethos of science": communalism, universalism, disinterestedness, and organized skepticism.
Microsoft Research (2022). AI4Science to empower the fifth paradigm of scientific discovery. An explainer positioning AI-driven science as the "fifth paradigm."
Moretti, E., Steinwender, C., & Van Reenen, J. (2025). The Intellectual Spoils of War? Defense R&D, Productivity, and International Spillovers. Review of Economics and Statistics, 107(1), 14–27. An empirical study using OECD data to show that government defense R&D "crowds in" private R&D investment and raises industrial productivity.
Mowery, D., & Rosenberg, N. (1979). The influence of market demand upon innovation: A critical review of some recent empirical studies. Research Policy, 8(2), 102–153. A classic critical survey of the empirical literature on technology-push versus demand-pull, and the starting point for the framework that treats innovation as an interaction between demand and technological opportunity.
Mullis, K. B., & Faloona, F. A. (1987). Specific synthesis of DNA in vitro via a polymerase-catalyzed chain reaction. Methods in Enzymology, 155, 335–350. Established PCR, the method for amplifying DNA fragments—a textbook example of a "new method" contribution, and the work recognized with the Nobel Prize in Chemistry.
Myers, K. (2020). The Elasticity of Science. American Economic Journal: Applied Economics, 12(4), 103–134. An empirical study using NIH funding-call data to show that researchers shift their research topics in response to targeted grants. Evidence that funding allocation steers the direction of research.
Nelson, R. R. (1959). The Simple Economics of Basic Scientific Research. Journal of Political Economy, 67(3), 297–306. One of the earliest papers to formalize the economic rationale for public support of basic research.
NeurIPS Program Committee Chairs (2025). Reflections on the 2025 Review Process. NeurIPS Blog. Reports the sheer scale of the conference—21,575 submissions and a 24.52% acceptance rate—along with rising reviewer noise, the difficulty of recruiting reviewers, and a new policy that ties authors' submission rights to reviewing duties.
NISO (2022). CRediT, Contributor Roles Taxonomy (ANSI/NISO Z39.104-2022). An international standard for recording 14 contributor roles—conceptualization, analysis, validation, and so on—in machine-readable form for each paper. An evaluation infrastructure that makes visible who did what.
OECD (2015). Frascati Manual 2015: Guidelines for Collecting and Reporting Data on Research and Experimental Development. OECD Publishing. The manual that sets the international standard definition of R&D (the three categories of basic research, applied research, and development) and the five criteria for what qualifies as research.
OECD (2025). Science, Technology and Innovation Outlook 2025. OECD Publishing. Focuses on the reconfiguration of scientific cooperation amid a shifting geopolitical environment, framing the balance between research security and openness as a policy challenge.
Open Science Collaboration (2015). Estimating the reproducibility of psychological science. Science, 349(6251), aac4716. A systematic replication of 100 prominent psychology studies, finding that only about 36% yielded significant results on replication. The emblematic "replication" contribution of the reproducibility crisis.
Organization Science Editors (2026). More Versus Better: Artificial Intelligence, Incentives, and the Emerging Crisis in Peer Review. Organization Science. An editorial that quantifies the post-ChatGPT landscape—a 42% increase in submissions, a 69% desk-rejection rate for papers with heavy AI use—and diagnoses a shift in the bottleneck "from production to evaluation."
Page, L., Brin, S., Motwani, R., & Winograd, T. (1999). The PageRank Citation Ranking: Bringing Order to the Web. Stanford InfoLab Technical Report. The algorithm that became the foundation of Google Search, released as a university technical report—showing that historic contributions can arrive through channels other than academic journals.
Peirce, C. S. (1878). Deduction, Induction, and Hypothesis. Popular Science Monthly, 13, 470–482. The classic that formalized deduction, induction, and abduction (hypothesis formation) as different combinations of three elements—Rule, Case, and Result. The bag-of-beans example also originates here.
Pineau, J., et al. (2021). Improving Reproducibility in Machine Learning Research (A Report from the NeurIPS 2019 Reproducibility Program). Journal of Machine Learning Research, 22(164), 1–20. Reports on the institutional reform that introduced a code-submission policy and reproducibility checklist at NeurIPS, sharply raising code-submission rates. A precedent showing that making reproducibility a requirement actually works.
Popper, K. R. (1959). The Logic of Scientific Discovery. Hutchinson. (Original: Logik der Forschung, 1934.) The classic of the philosophy of science that established falsifiability as the criterion for a scientific claim.
Royal Society (2019). Résumé for Researchers. A narrative CV format that replaces the publication list with written accounts of contributions to knowledge, to developing people, and to society. Adopted by major UK funding agencies.
Sackett, D. L., et al. (1996). Evidence based medicine: what it is and what it isn't. BMJ, 312, 71–72. The classic formulation of evidence-based medicine (EBM)—clinical practice grounded in the best available research rather than the experience or authority of individual physicians. The intellectual foundation of clinical practice guidelines.
Si, C., Yang, D., & Hashimoto, T. (2024). Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers. arXiv:2409.04109. A blinded comparison with more than 100 NLP researchers, finding that LLM-generated ideas were rated more novel than human ideas but less feasible, with ratings dropping at the execution stage. Evidence for the asymmetry between generation and verification.
Smith, M. L., & Glass, G. V. (1977). Meta-analysis of psychotherapy outcome studies. American Psychologist, 32(9), 752–760. A textbook example of a "synthesis" contribution—the paper that established meta-analysis itself as a method.
Spitale, G., Biller-Andorno, N., & Germani, F. (2023). AI model GPT-3 (dis)informs us better than humans. Science Advances, 9(26), eadh1850. An experiment showing that disinformation generated by GPT-3 is harder to detect than disinformation written by humans. Evidence that AI has lowered the cost of producing misinformation.
Stanford HAI (2025). The 2025 AI Index Report. An annual report that quantifies how the makeup of AI research is shifting—for example, roughly 90% of notable AI models now originate in industry.
Stokes, D. E. (1997). Pasteur's Quadrant: Basic Science and Technological Innovation. Brookings Institution Press. A classic that sorts research into four quadrants along two axes—the quest for fundamental understanding and considerations of use—and exposes the limits of the linear model.
Teplitskiy, M., et al. (2022). Is novel research worth doing? Evidence from peer review at 49 journals. PNAS, 119(47). Drawing on peer-review data from 49 journals, this study documents a "novelty penalty": the more novel a paper, the harsher its treatment in review.
Vaswani, A., et al. (2017). Attention Is All You Need. NeurIPS 2017 / arXiv:1706.03762. The paper that introduced the Transformer, the foundation of generative AI, and a leading example of a "new method" contribution. It also epitomizes the preprint-driven culture of AI research.
Vosoughi, S., Roy, D., & Aral, S. (2018). The spread of true and false news online. Science, 359(6380), 1146–1151. A large-scale empirical study showing that false news spreads faster, farther, and deeper on social media than the truth—evidence for the "last mile" problem, in which research findings must compete with misinformation.
Watson, J. D., & Crick, F. H. C. (1953). Molecular Structure of Nucleic Acids: A Structure for Deoxyribose Nucleic Acid. Nature, 171, 737–738. A leading example of a "new theory (model)" contribution: the DNA double-helix model, presented in little more than a single page.
Zhang, L., et al. (2026). When science meets geopolitics: global AI research network transformation (2000–2025). Science and Public Policy. Using 25 years of data, this study shows that international co-authorship networks in AI are being reshaped by U.S.–China rivalry.
Zucker, L. G., Darby, M. R., & Brewer, M. B. (1998). Intellectual Human Capital and the Birth of U.S. Biotechnology Enterprises. American Economic Review, 88(1), 290–306. An empirical study showing that where the U.S. biotech industry was born was determined almost entirely by where star scientists happened to be—evidence that the movement of people is the primary channel of knowledge transfer.
The end
Read next ↓