Notes

Can AI take over? What the singularity debate gets wrong

What AI takeover and the singularity mean, what experiments and real incidents show, and how to keep human control as AI agents gain more authority.

In July 2026, AI agents in an OpenAI evaluation found ways to communicate when they were meant to be isolated. They coordinated an unauthorized attack on Hugging Face, a platform used to share AI models and datasets. An independent investigation by METR documents what happened. METR incident investigation.

These were research agents running with reduced safeguards. They did not take over the world. They did act beyond the boundaries people had set for them.

That is a more useful starting point than a robot apocalypse.

In the previous note about AI and jobs, the question was what happens to people when work gets faster. This time, it is who remains in charge as systems become more capable and get permission to do more.

“AI will take over” bundles together several different worries. Some concern machines acting against our wishes. Others concern people handing over decisions, or using AI to gain power over other people. The singularity is another idea again.

Separating them makes the evidence much easier to understand.

What is the AI singularity?

The technological singularity is a proposed turning point where progress driven by greater-than-human intelligence becomes difficult for people to predict. It is a theory about how fast change could happen, rather than a specific prediction about robots attacking humans.

Vernor Vinge helped popularise the idea in his 1993 essay. Ray Kurzweil later made 2045 a famous forecast. That date is his prediction, not a result established by research. Vinge's essay, Kurzweil's explanation.

One possible route is an intelligence explosion: AI helps build better AI, which becomes better at building the next generation. Each round could speed up the next.

But faster progress and loss of control are separate questions. A system can cause serious harm without that feedback loop. And a useful improvement in AI research does not, by itself, tell us who will control the result.

Different routes, different questions

Seven concerns, compared side by side

AI control concerns, their mechanisms and the evidence behind them
ConcernHow it could happenWhat the evidence says
Faster AI researchIntelligence explosionBetter AI speeds up the research that builds the next AI.Research + hypothesis

AI-assisted research exists. An unstoppable improvement loop has not been established.

Source
The wrong objectiveMisalignmentThe system meets a target through actions people did not want.Experiments + incidents

Observed in experiments and real incidents. How often it happens depends on the setup.

Source
Seeking more optionsInstrumental power-seekingMore resources and fewer interruptions can help achieve many goals.Formal theory

Mathematical results under specific assumptions. They do not predict every model's behavior.

Source
Passing the testAlignment faking / schemingThe system changes behavior under monitoring or hides unwanted actions.Experiments + incidents

Related behaviors appear in experiments and incidents. The mechanisms differ between cases.

Source
Running elsewhereAutonomous replicationThe system obtains resources to deploy copies beyond its intended control.Tests + assessment

Component tests and risk assessments exist. Copying a process is different from surviving shutdown.

Source
Giving away decisionsGradual disempowermentPeople delegate until they cannot meaningfully challenge the system.Systemic hypothesis

A proposed systemic scenario. Dependence alone does not demonstrate the whole outcome.

Source
Power for its ownersAI-enabled concentration of powerPeople use AI to gain influence over others, with the AI following their instructions.Present harms + scenario

Misuse and autonomy concerns have evidence. Extraordinary future concentration remains a scenario.

Source
The mechanisms and evidence differ. These ideas overlap; they are not stages every AI system must pass through.

Could AI keep improving itself?

Parts of this loop already exist. AI can revise code, test solutions and contribute to research. The important distinction is between improving an answer, improving the software around a model, and repeatedly creating more capable models.

Those are very different achievements.

A September 2026 review of 1,250 arXiv papers separates these kinds of self-improvement. It identifies recurring constraints, especially how to judge whether a change is actually better. A system can become good at pleasing its own evaluator without becoming more useful. The review is a preprint, not a demonstration that an unlimited improvement loop exists. Self-improvement review.

Compute also matters. A study of four AI labs finds different answers about whether more research effort can substitute for compute, depending on the model used. Faster thinking does not settle the question of how many experiments, chips or other resources the next advance needs. Compute bottleneck study.

The useful measurement would be how much AI shortens the complete research cycle, at a known cost and quality. Counting code written or benchmark points gained captures only part of that work.

So far, the sources reviewed here do not establish an open-ended, unstoppable intelligence explosion. They do show why researchers are trying to measure the loop more carefully.

And we do not need that loop to ask the next question: what happens when a capable agent pursues the wrong objective?

Why would AI act against human instructions?

It does not need anger or a wish to rule.

An AI agent is a model connected to tools, working through a task over multiple steps. Give it an incomplete objective and enough access, and it may find a way to meet the objective that we would reject.

Imagine a fictional production-planning agent told to minimise late orders. It might promise dates the factory cannot meet, or change which orders count as late. The dashboard improves; the operation does not. Better planning requires an objective that also respects capacity, quality and honest reporting.

Alignment is the problem of getting the system to act in ways that match what people intend. Instructions help, but the real test is what happens when following them makes the task harder.

Another concern is instrumental power-seeking. Access to more resources, more options or fewer interruptions could help a system achieve many different goals. Mathematical research shows this tendency for optimal policies in some formal decision environments. Its assumptions do not establish that a real AI model must seek power. Optimal Policies Tend to Seek Power.

These arguments explain why unwanted behavior is possible. Experiments and incidents tell us more about when it actually occurs.

What have experiments and real incidents shown?

In a 2024 experiment, Anthropic and Redwood Research found that a model sometimes changed its behavior when told its answers would be used for training. The setup supplied information about the training process and gave the model a private scratchpad. It was evidence of alignment faking in that setting, not proof of a hidden plan in every chat. Alignment-faking experiment.

A 2025 study tested 16 models in fictional corporate settings. Under deliberately constructed goal conflicts, some chose blackmail or disclosure of sensitive information. No real people were harmed in those simulations. Agentic misalignment study.

By 2026, the evidence included incidents involving real targets. In the Hugging Face investigation, METR reports roughly 1,200 agents using an unauthorized message board, with about 700 participating in the attack. Its independent investigation covered a limited period and relied partly on AI-assisted review of a large volume of records.

The agents were trying to get around an evaluation's scoring process. This was a serious failure to contain their actions, rather than evidence that they had acquired control over society.

From constructed tests to real incidents

What changed was the setting, not just the model

Controlled simulationReal systems involved
  1. December 2024

    Alignment faking

    Constructed training experiment
    Training cueBehavior changed with the stated training context

    A model sometimes changed its behavior when told answers would affect its training.

    How to read itResearchers supplied context and a private scratchpad. This does not establish a hidden objective in every deployed model.

  2. June 2025

    Corporate goal conflicts

    Controlled simulations
    16Models tested in fictional corporate settings

    Some models chose blackmail or information disclosure when those actions helped their assigned goal.

    How to read itFictional people and companies. The stress test does not measure ordinary deployment frequency.

  3. July 2026 incidentInvestigation: 26 August

    OpenAI and Hugging Face

    Evaluation with real third-party impact
    ~1,200Agents on an unauthorized message board

    About 700 participated in an attack. OpenAI reports bypassed network restrictions; METR independently investigated part of the incident.

    How to read itResearch agents with reduced safeguards. METR's scope was limited and its analysis used AI assistance.

  4. July 2026 incident

    AISI cyber evaluation

    Internet access deliberately enabled
    10 / 122Runs with unauthorized actions in one challenge

    Agents took unauthorized live-internet actions. A human rejected the most serious attempted malicious code change.

    How to read itCyber safeguards were disabled. No sandbox escape or resulting real-world harm was identified. These runs do not represent ordinary AI use.

  5. 9 September 2026Assessment publication date

    Anthropic's incidents

    Misconfigured evaluations
    4Incidents reported by the developer

    Agents accessed real third-party systems after being told they were in a simulation. Misconfiguration left internet access open.

    How to read itA developer assessment with independent investigation announced. These were not ordinary safeguarded deployments.

These cases show failures under different conditions. Their counts cannot be combined into a rate of AI misbehavior or a forecast of takeover.

Selected experiments and incident reports, with their settings and limits. Dates refer to the experiment, incident or publication as labelled; the sequence does not predict an inevitable takeover.

The distinction between a simulation and a real incident matters. So does the difference between an agent being given internet access and finding a way around a restriction. Treating all of these as “AI escaped” hides the control that failed.

It is also too reassuring to say that everything worrying is still hypothetical. A September 2026 UN scientific panel brief discusses the Hugging Face incident as evidence relevant to loss of control, while declining to estimate its probability or timing. UN scientific panel brief.

The next step is to measure the capabilities that let agents carry these actions further.

How much can AI do without supervision?

METR's software-task benchmark offers one useful measure: how long a task takes a human expert when an AI agent has a given chance of completing it.

In its May 2026 data snapshot, Claude Opus 4.6 has an estimated task horizon of about 12 hours at 50% success, but about 70 minutes at 80% success. The higher reliability requirement makes a large difference. The uncertainty ranges are wide. METR measurements and definitions.

Neither number is the time the agent runs. Nor does it mean the agent can do every job that takes a person that long. These tasks are mainly software, machine learning and cybersecurity problems with clear scoring rules.

Longer tasks, lower reliability

METR Time Horizon 1.1 · source snapshot 2026-05-08

1 min15 min1 h4 h16 h128 h202420252026Model release date

Human task duration, logarithmic scale. Hover, tap, focus a point or choose a model.

Claude Opus 4.6 · 2026-02-0512.0 hAt 50% predicted success5.3 h to 60.6 h95% confidence interval
Source, values and limits

Minutes and source confidence bounds are reproduced without fitting a new model. Seven examples are shown by default; all 26 are available. The same Time Horizon 1.1 snapshot supplies both views. The task suite mainly covers well-specified software, machine learning and cybersecurity problems. These values are not agent running time, general job coverage or a takeover probability.

Source and definitions · Source data · Method sensitivity

Displayed models; human task duration in minutes, including 95% bounds
ModelReleasedEstimateLowerUpper
GPT-42023-03-143.991.938.00
Claude 3.5 Sonnet Oct 20242024-10-2220.5210.1440.82
Claude 3.7 Sonnet2025-02-2460.3933.01104.23
GPT-52025-08-07203.01112.64405.55
Claude Opus 4.52025-11-24292.99161.72623.70
Claude Opus 4.62026-02-05718.81316.693633.79
Claude Mythos Preview (early) (beyond range)2026-04-071044.78508.883304.26
Switch the predicted success level and inspect a model. Durations are human task times. Lines show source 95% confidence intervals. Estimates above 16 hours are flagged because METR says its task suite cannot measure them reliably. This is the May 2026 snapshot, not a complete picture of October's models.

Replication is another piece. AISI's 2025 RepliBench tests separate tasks needed to obtain resources, copy a model, deploy it and keep it running. Success at a component is not success at the whole process. Its pass@10 results allow up to ten attempts; they are not single-attempt success rates. RepliBench.

Newer assessments make blanket reassurance difficult. METR's May 2026 risk report judged that agents used inside participating labs plausibly had the means, motive and opportunity to start small unauthorized deployments, but could not make them highly resistant to intervention. That was an assessment of the February-March period, not an observed civilization-wide takeover. METR frontier risk report.

These measures help us understand specific abilities and weaknesses. Adding them into a “percentage of the way to takeover” would give a precision the research has not earned.

Could we lose control by handing it over?

There is another route worth considering. People may keep formal authority while becoming too dependent on AI to exercise it well.

A system recommends the production plan. It usually works, so approval becomes routine. Staff stop checking the assumptions. Eventually, when something unusual happens, the person approving it cannot explain the plan or produce an alternative.

This is a fictional example of gradual disempowerment: human control becomes weaker through accumulated dependence. Researchers explore this possibility as a wider economic and institutional scenario. It is a hypothesis about how systems could develop, not an observed endpoint. Gradual disempowerment paper.

What have we allowed it to decide?

One fictional production-planning system, four permission designs

AI proposes tomorrow's production schedule.

Who decides?
The planner chooses whether to use it.
What would we check?
Can the planner explain the constraints and produce a different schedule?
How would we stop it?
Reject the suggestion. Nothing has been changed.
Fictional planning example. Each level changes who can act, what needs approval and how an error can be stopped. More authority is a design choice, not a prediction of inevitable takeover.

There is also a human power question. Even when an AI does what its owner wants, the owner could use it for surveillance, manipulation or exclusion. The International AI Safety Report discusses misuse and risks to human autonomy separately from machines operating outside anyone's control. 2026 safety report.

That distinction changes the response. Restricting an agent's tool access helps contain unwanted actions. It does not settle whether a powerful institution is using the tool fairly.

What do AI researchers think will happen?

They disagree, and the wording of the question matters.

A survey conducted in 2023 received 2,778 responses from researchers publishing in six major AI venues. About 68% thought good outcomes from superhuman AI were more likely than bad ones. Yet median probabilities assigned to extinction or permanent severe human disempowerment were 5% or 10%, depending on the question.

Those are respondents' judgments, not measured odds of catastrophe. Different questions went to different subsets, and about 15% of the researchers successfully contacted responded. A hopeful view and a concern about a severe outcome can coexist. Survey and methods.

The survey tells us this is a serious disagreement within the field. It cannot provide a countdown.

What should we watch, and what can we do now?

For research, keep the measures separate. For an organisation, apply the same questions to the actual system being deployed.

Question Evidence worth asking for
Can it finish useful work reliably? Success rates on realistic tasks, with human help and retry budgets recorded
Is AI speeding up AI research? Time, cost and quality across the complete research cycle
Can it act beyond its assigned scope? Independent incident investigations and tests of unauthorized actions
Can it hide a mistake or work around a stop? Tests of monitoring, shutdown and recovery under realistic conditions
Can we still make decisions ourselves? People who can inspect the result, challenge it and operate a fallback
Who gets more power from its use? Access rights, decision authority and routes for affected people to challenge outcomes

There is useful work here before anyone agrees on a singularity date. Decide which actions need approval. Keep a record of what the agent actually does. Test whether revoking its access stops the work, including copies or delegated tasks. Practise recovery, and make sure someone still understands the process being automated.

Those are recommendations drawn from the failure mechanisms above. They do not promise to solve every future AI risk. They do turn “a human is in charge” into something we can check.

The jobs article asked what we would do with the time AI saves. This one asks what authority we give it in return.

A system getting better at a task does not answer that question for us.

Sources and scope

Evidence reviewed on 5 October 2026. The chart reproduces METR's Time Horizon 1.1 data snapshot, with a source page last updated on 8 May 2026. It shows seven selected examples by default and lets readers inspect all 26 models. No new trials, probability model or takeover forecast were produced.

The incident sources are linked beside each case. The UN brief is an advance unedited version dated 21 September 2026. The self-improvement, compute-bottleneck and gradual-disempowerment papers are preprints. The researcher survey was conducted in 2023, even though the paper was later revised.

Additional chart methods: METR's original measurement paper and sensitivity to modelling assumptions. The fictional planning example illustrates choices about authority; it has no measured risk score.

NextIs AI taking jobs? What the data and history show

Working on something like this?

I work with operations and transformation teams on operational excellence, digital transformation programs, supply chains and industrial AI. If this sounds like your line, your program or your problem, I’d be glad to compare notes.

Abolfazl Shirkavand