OpenAI says coding agents have reached research intern level, targets automated researcher by 2028
- OpenAI says it has reached its September goal for an automated research intern, a system that completes well-defined research tasks under human direction that would take a skilled researcher several days.
- The company is targeting an automated AI researcher by March 2028 to advance deep learning and alignment under human supervision, while people retain control of priorities, results, scaling, pauses, and deployment.
- By mid-August, the median OpenAI researcher used coding agents daily at more than $600 per day in API-priced inference, with agents taking on more complex tasks and succeeding more often.
- After the recent Hugging Face incident, OpenAI paused reinforcement-learning training for its latest deployment-bound models, hardened and red-teamed research environments, and resumed only some workloads under stricter controls.
- OpenAI says it cannot assume alignment and safety will keep pace with capabilities or that stronger systems will remain monitorable, and says it will slow or stop development or deployment if it cannot sufficiently safeguard systems.
Hacker News opinions
The opening is hard to get through, but the internal details about researchers using OpenAI's tools are interesting. I kept looking for a definition of RSI and never found one. That acronym is not common outside OpenAI's circle.
I see RSI as starting with human tool use. You could argue it begins whenever intelligence improves its ability to act in a physical environment.
RSI sounds like singularity jargon. LLM capability depends on data and compute, so automating experiments or producing synthetic data does not create unlimited compute for testing ideas at scale.
I think OpenAI is trying to normalize RSI as routine and safe before the term becomes familiar outside AI safety circles. The posts feel like marketing mixed with an attempt to frame the debate early.
What I want to know is whether OpenAI would roll back to a known-safe checkpoint if an earlier misaligned model contaminated later generations. I doubt it would unless someone forced the issue.
I do not think anyone can know that a model has transmitted no tendency toward misalignment. Using models with known minor misalignment to train stronger ones, then using the same model lineage for safety and steering, seems risky.
This feels close to an impossible mission. How do you perfectly control and observe a human-level mind, and how deterministic is a rollback anyway?
My experience roughly matches the automation claim: stronger models and better tooling let me run unattended jobs 24/7 starting in March, using an Anthropic subscription and my own hardware. OpenAI's claimed $8,000 per researcher per day is still wild, and I want to know how they track the work.
OpenAI may be optimizing for token use, so an employee could spend $8,000 generating an animated pelican and still climb an internal leaderboard. I wonder whether someone spending roughly $300,000 to translate the FLT proof into Lean gets rewarded differently.