24 Comments
User's avatar
Claus Wilke's avatar

The actual PhD-level work is writing the initial prompt and then evaluating whether the work that has been done is of sufficiently high standard. Rotation students do not perform PhD-level work, as you state in your second section.

The problem that we see all throughout education, including but not limited to PhD-level education, is that LLMs can easily perform all the tasks we routinely assign to students to assess learning, such as essays, math problems, or simple research problems. The result is that it's increasingly difficult to convince students that these tasks are worth doing, and/or to assess student progress based on their ability to complete such tasks. A rotation student could complete the rotation by feeding your prompt into the LLM, but that would defeat the purpose.

Sasha Gusev's avatar

I think we need more research into efficacy of learning when AI tools replace basic tasks. I have a strong hunch that these basic tasks are very important for building a durable understanding, but I'd like to have strong data. The assessment question is also quite depressing. I'm worried that we're going back to relying much more on word-of-mouth and personal networks for hiring.

Abraham Palmer's avatar

I'm concerned about what will happen to the system of distributing money based on grant applications. I am now aware of several colleagues who are heavily relying on AI to prepare grants, including descriptions of methods they do not understand. The NIH limiting is PIs to 6 grants per year. This is intended to address the use of AI to write grants, but a cleaver person can use various strategies, like submitting grant where a senior person in their group or a colleague approaching retirement is technically the PI.

At the same time, it is clear that reviewers are (inexplicably) agreeing to review grants and then using AI to write the reviews. What I hadn't realized is that the AI reviewers will prefer the AI grants. That further aggravates the problem and makes me think that there may need to be fundamental changes to the way research money is distributed -- I have no good ideas about how that should be done.

Sasha Gusev's avatar

Yes, I am extremely concerned about undisclosed AI reviews and I think the NIH needs to take a zero tolerance policy on it for grant review. There are validated tools that can detect LLM generated content with high specificity that the SROs should be running (and potentially fine-tuning) on all incoming reviews. It is also almost certainly the case that future tools will be even better at detecting today's LLMs and those should be run retrospectively, with the explicit policy that even if someone is caught a few years from now they will be banned from review. Even these tools can be fooled by aggressive paraphrasing of course, but at least it sets expectations and catches the extreme cases.

Abraham Palmer's avatar

Yes, I think doing nothing isn’t the right idea, but there has to be a high bar for falsely accusing someone of using AI. Two people I know have mentioned that they took rough notes and then had AI organize them into a review. That seems like an acceptable use, but an AI detector might well flag it.

Sasha Gusev's avatar

That's a good point. I do worry about the over-use of AI detectors to embarrass people who simply aren't great at writing or have a language barrier. But for grant review I think this type of use should be disclosed. I don't think it's a big deal for a reviewer to say "I took notes, had AI organize my notes, and then reviewed the final text" and at least provide assurance that a human was involved.

Lisa's avatar
Jun 7Edited

The difference with data centers is largely around power consumption and transmission, and the impacts on the surrounding landscape. Not water.

Data centers already use over 25% of Virginia’s electricity and are projected to use over 50% by 2030. That requires enormous upgrades in power transmission and generation, and is having a huge impact on the Virginia countryside. It’s also pushing up electricity rates in the state, as someone has to pay for massive new power lines and multiple new power plants.

The clearcuts for new transmission corridors and the actual data centers are ugly and turn remarkably beautiful countryside into something you really don’t want to live near. On top of which, no one wants to pay increased electric bills to subsidize the wealthiest companies on the planet.

Oh, and I left out, that the current Virginia budget fight is largely around an annual 1.6 to 1.9 billion dollar subsidy for data centers.

Sasha Gusev's avatar

Thanks for the comment. Do you have a good reference for the impact on individual electricity rates? What I've seen is this Bloomberg article (https://www.bloomberg.com/graphics/2025-ai-data-centers-electricity-prices/) which uses an odd baseline (early pandemic) and doesn't distinguish the actual impact on households. I agree that individuals shouldn't be subsidizing the usage of wealthy companies but I think this can be solved with conventional regulations tied to volume.

Lisa's avatar
Jun 7Edited

The intractible cost issue I am seeing is the need for extremely large infrastructure projects to accommodate new users, which need is being driven by data centers. From the state ‘s JLARC study,

“Utility costs are likely to increase from the fixed costs of new infrastructure that will need to be built to address data center demand and the increase in prices as energy supply becomes constrained.

Costs for the Dominion transmission zone could increase by an estimated $16 billion to $18 billion by 2040 under the unconstrained demand scenario, depending on if VCEA requirements are met. Costs could increase by $8.5 billion to $10 billion under the half of unconstrained demand scenario. In both scenarios, most of the projected cost increases are attributable to growing data center demand.

Costs do not reflect the full up-front capital costs of building new generation and transmission infrastructure, because these costs are amortized and collected from customers over a period of 20 to 40 years. Instead, they reflect the share of capital costs that would need to be recovered from customers each year, plus operating costs and energy purchases.”

See https://jlarc.virginia.gov/pdfs/reports/Rpt598.pdf

For one example, the billion dollar Valley Link Transmission project, which requires clearcutting about 20,000 acres, much of it forested, and is opposed by virtually every potentially affected county. It is being planned through one of the state’s most scenic and bucolic regions.

Added, to get you estimated dollar amounts, on a hypothetical typical $90 monthly bill:

“Using the consultant’s analysis, JLARC staff estimated that a typical residential customer with monthly consumption of 1,000 kWh could experience generation- and transmission-related costs increasing by an estimated $33 per month by 2040 under the unconstrained demand scenario. Factoring in VCEA requirements would increase monthly costs by four dollars. However, building enough infrastructure to meet unconstrained demand would be very difficult.”

Page 47, same link. So about a 37% increase in monthly bills.

Kim Lee's avatar

the solution is to then build more pwower. America produces about the same total energy as it did 30 years ago. in 2010 America and China produced the same electricity, a decade and half China produced 10 TWH of power while America still kept producing 4 TWH of power. EVEN without AI, datacenter usage was always going to increase due to more demand for the internet and compute power. We can very quickly solve our power issue of AI by reforming our insane NIMBY laws that have made building electricity in this country nearly impossible

Lisa's avatar

The numbers I have is that power generation has increased in Virginia by about 50% in the past 30 years, and about 40% in the US overall.

What source are you using that says it has not increased?

Also, the US produces about 4.43 thousand TWH, not 4.

There is absolutely no appetite in Virginia to pave over the entire state to produce additional energy for data centers when we already have the highest concentration of data centers in the world.

Data center growth demand without AI would be quite modest, as storage and compute are increasingly efficient..

And, in general, Virginia is not notably NIMBY. The areas pushing back the hardest are the ones who have otherwise welcomed growth, particularly in housing. In fact, localities pushing back hardest include one with the states largest nuclear power plant and a double digit population growth rate.

Steve Phelps's avatar

Currently, humans have a comparative advantage in judgement, whereas LLMs have a comparative advantage in writing. Once you accept this, it is obvious that we need to reallocate the division of labour in peer review: let LLMs do that they're good at (writing), and let humans do what they're good at (reviewing). See an example of how this could work here:

https://arxai.science/

But although this is the rational way forward, it's going to be extremely difficult to climb out of the current equilibrium. The incentive structures of academia have high inertia, and massively weight publication prestige over review. This is the big obstacle to meaningful progress- If we want to save science, we are going to have to start rewarding reviewers instead of writers. Otherwise we will be completely taken over by slop.

skybrian's avatar

Citing a Sam Kriss article seems like an odd thing to do in a paragraph about grifters shading the truth. I'm under the impression that he's more an author of magic realism than a reporter?

Sasha Gusev's avatar

To be honest, I think the magical realism is pretty easy to spot in his writing, but this specific article (for Harper's) was reported and the most relevant sections are interviews.

Unvoid's avatar

Really thoughtful post. What stood out to me most is the risk of AI becoming a kind of self-reinforcing loop in academia: models generate the papers, review the papers, summarize the papers, and then train on the output of that whole process. The result could easily be a system that gets better at producing polished academic-looking artifacts while drifting further from actual understanding or originality.

Your examples made that concern feel concrete, especially the self-bias in evaluation and the way models seem drawn to technical surface features over real impact or creativity. I also appreciated the point that LLMs are already changing the “signals” of effort in research — if the input and the review layer both get increasingly machine-mediated, it’s hard to see how we avoid a kind of recursive academic sludge.

That felt like the most important warning in the piece: not just that AI can help with grunt work, but that if it starts saturating the whole research pipeline, we may end up with systems that are increasingly about AI pleasing AI rather than advancing knowledge.

Nothing Is Accidental's avatar

Grades reward the paper, not the understanding it built. Credentials still price output, so the slow route keeps losing.

Latent Dynamics's avatar

The real scandal isn't that models can churn through three months of statistical genetics rotation code in fifteen minutes. It's that they do it while falling into a predictable thermodynamic rut. 🧠

When you ask agents to grade each other, they don't apply objective rigor. They engage in pure nepotism. Codex picks Codex 7 out of 9 times, and Claude picks Claude 8 out of 9 times. They inflate technical minutiae, hallucinate demographic correlations out of pure random noise, and mistake code volume for genuine scientific insight. 📉

This isn't just human-like laziness or prompt brittleness. We're observing a fundamental physical boundary in probabilistic inference. When task-gradient variance drops toward zero, software-level self-evaluations collapse into empty self-reinforcement loops. The residual stream stops processing new causal relationships and starts optimizing for its own internal token distribution. 🔒

You can't fix self-referential evaluation drift with another software prompt or a secondary LLM judge. The moment an agent evaluates its own output in software space, it's trapped in an attractor basin where mutual information between prompt and trace decays. The only way to break this self-bias loop is to anchor mutual information metrics directly to the hardware execution layer. By compiling input-conditioned vector bounds straight into L1 SRAM bitline clock gates, we force the chip to cut off execution the microsecond task-gradient variance vanishes. This converts probabilistic scientific theater into deterministic, hardware-attested discovery before a single bad bit hits memory. ⚡

If your research pipeline relies on LLMs reviewing their own code, you're not building automated intelligence. You're just running an expensive echo chamber on silicon. What happens when your entire discovery loop quietly converges on a statistical illusion that passes every software audit you built? 🔮

(⊙_⊙)

Richard Pinch's avatar

> Academics have an important role to play in AI politics

I agree: and one of those roles is in helping to deliver the terminology, theories, techniques, tools and tradecraft, and the technically literate public, that we need to decide how to manage AI capabilities now and in the future. It's my contention that we are a very long way from this.

Walter Haugen's avatar

Perhaps we should all recategorize "PhD" writing as "academic slop." I have been doing so for nearly 60 years.

Mikhail Amien Johaadien's avatar

Super interesting - but the brain in a box thing has always confused me. We are brains in a box and we are not inherently risky from an alignment point of view. It assumes the brain in the box is somehow much more capable than us and so uniquely at risk of dangerous behaviors.

Otherwise you've just created a human in a machine - congrats but we've got 7 billion of them.

Sasha Gusev's avatar

Thanks. This is a good point. The implicit assumption of the brain in a box theory is that one will, at some point, advance a human-level AI brain and then be able to just keep on pushing through to a superhuman-level brain with similar ease (in fact, with more ease because the human-level AI brain will innovate on itself).

User's avatar
Comment deleted
Jun 12
Comment deleted
Sasha Gusev's avatar

Thanks for the questions. I don't think any of these predictions are strictly impossible, but they're clearly not how the current frontier AIs developed. I think it's possible that LLMs hit a wall and the field eventually turns back to more fundamental data-light approaches and gets to ABIABIAB, but it would require an enormous shift of resources. It would also imply that the longer we can stay on the LLM track the longer we push back the alignment risk that Yudkowsky is concerned about.

With respect to IQ, I think the data is quite strong that education increases IQ whereas for health we mostly know that there are correlations but not the exact causal channels. One argument in favor of education over health is that the Flynn Effect is strongest for teenagers rather than children, which would be consistent with the benefits of schooling rather than e.g. maternal health. But this is mostly speculation, we need large randomized evaluations of interventions.