
The Debate / Analysis
When AI Researchers Warn About Superintelligence —
What Should We Actually Do?
The case for building intelligence with humanity, not beyond humanity.

Start by taking the warning seriously.
A warning about superintelligence deserves more than a reassuring slogan. It also deserves more than a headline that treats a possible future as an established fact. The useful question is what evidence, institutions and engineering choices would make continued development defensible.
In his September 2026 resignation thread, Jacob Coxon argued that frontier laboratories are racing toward self-improving superintelligence without acting responsibly. He described competition as an obstacle to caution, called for coordination, and said preventing a global race might require a temporary ban on capability improvements. His argument is not simply that powerful AI is frightening: it is that proceeding under these conditions requires an extraordinary level of justification.
Here is our strongest reconstruction of the underlying argument: if a development process could produce systems able to defeat their own oversight, and the available safeguards cannot be tested with sufficient confidence, waiting for widespread damage would be an unacceptable way to learn. Competitive pressure can make every participant feel that slowing down alone would simply hand the future to someone else. Individually understandable choices could still produce a collectively dangerous outcome.
That argument does not require hostility to intelligence. It asks who bears the risk, who can refuse it, and whether the people taking the decision can also prevent the consequences. Answering it requires evidence and governance, not a declaration that our intentions are good.
The July 2026 Pacing the Frontier statement makes a related institutional argument: develop international tools for deliberately pacing automated AI development when necessary. It is an advocacy position from its signatories, not a measured estimate of catastrophe.
The strongest case for continued development.
A serious counterargument begins with the costs of withholding useful capabilities. Better tools could support scientific work, accessibility and difficult information tasks. Restrictions that do not distinguish a narrowly scoped application from a frontier training programme could block benefits without addressing the mechanism of danger. This is an argument about policy design, not a guarantee that every AI deployment is beneficial.
There is also an engineering case for learning through carefully bounded experiments. Some control measures can be tested before giving a system consequential authority. The cited AI Control study evaluated protocols against deliberate subversion in a programming-task setting. Its results support investigation of concrete safeguards; they do not establish that the same protocols can contain arbitrary future superintelligence.
The limitation of this counterargument is equally important: useful research does not justify every scale, access level or release decision. A small experiment cannot silently become an unrestricted deployment. And the possibility that another actor might move first is not itself evidence that a proposed system is safe.
These are not two indivisible camps. A person can support AI research, favour restrictive rules for particularly dangerous capabilities, and oppose a blanket halt to low-risk work. The practical dispute concerns thresholds, verification and the authority to act when those thresholds are crossed.
What the evidence establishes — and what it does not.
The International AI Safety Report 2026, published in February, describes wide disagreement about the probability of severe loss of control. Its assessment distinguishes a system’s capabilities, its tendency to use them harmfully and the opportunities provided by deployment. It reported early signs of relevant abilities, but not the combined, sustained capabilities needed for the severe scenarios it analysed. That is a dated scientific assessment, not a verdict on every system available in September.
A failed safeguard in a test is evidence about that test and its threat model. It may reveal a serious weakness. It does not, by itself, measure the probability of a civilisation-wide outcome. Conversely, passing a benchmark is not proof that a system will remain controllable in an unfamiliar environment.
Describe the observed behaviour, setting, tools and limits of a result.
State assumptions about future capability, exposure and countermeasures.
Explain which risks are acceptable, to whom, and who gets to decide.
Open questions include whether oversight can remain effective as tasks become more complex, whether models behave differently under evaluation, and whether a reviewer has enough information and time to detect consequential errors. A personal probability estimate is not a laboratory measurement. Uncertainty should change the conditions of a decision; it should not be used to claim either certainty of disaster or certainty of safety.
Self-improvement also needs a precise definition. A system drafting code, running an authorised experiment and having its results reviewed is a different arrangement from a system acquiring resources and changing its own constraints without approval. Any argument about a runaway process should explain the feedback loop, the resources it depends on and the points at which intervention might still work.
What should we actually do?
IGIQANT’s editorial recommendation is to connect each increase in capability or autonomy to a specific burden of evidence. The following questions are proposed decision criteria, not a certification scheme.
Define the system and its exposure.
Specify the task, users, affected people, tools, data, resources and external actions. Identify who can be harmed even if they never choose to use the system. A capability score alone does not describe this exposure.
Test a credible failure story.
Write down how the system could act incorrectly, hide a problem or receive misleading instructions. Evaluate against that story and simpler baselines. Record failures and uncertainty alongside successful demonstrations.
Separate the ability to propose from authority to act.
Make access limits enforceable by the surrounding system. A prompt telling an agent to be careful is not an access-control boundary. Consequential actions need an explicit owner, an appropriate review path and a way to stop further action.
Make review workable.
Give reviewers sources, material assumptions and a concise account of the proposed action. Measure whether they actually detect errors. If the workload makes informed review unrealistic, reduce the scope or rate of action.
Set conditions for withholding, pausing and resuming.
Decide in advance which failures stop a trial, who has authority to stop it, and what evidence is required to restart. Where credible severe risks cannot be managed, withholding deployment or pausing a risky capability path must remain an option.
Keep responsibility outside the model.
Assign accountable people and organisations. Preserve incident records and challenge routes. Include affected communities in decisions about acceptable use; developer preference is not a substitute for their interests.
These recommendations draw on NIST’s risk-management functions and OWASP’s agency guidance. NIST organises risk work around governance, context, measurement and management; OWASP emphasises limiting functionality, permissions and autonomy. Neither framework is proof that an arbitrary future system is safe.
Public rules matter alongside engineering. The EU AI Act overview describes a risk-based framework, including requirements for relevant high-risk systems and general-purpose AI. Legal duties depend on the system and the actor’s role. Compliance and technical safety address overlapping questions; one should not be presented as a complete answer to the other.
The IGI Perspective
IGIQANT proposes a direction we call IGI — Infinite Growing Intelligence: intelligence designed to grow with humanity, not away from it. This is a design commitment to investigate. It is not a claim that we have created AGI, ASI, consciousness or a solution to alignment.
The expression AI → AGI → ASI → IGI can introduce a conversation, but it can also mislead. AI, AGI and ASI usually describe a field or classes of capability. Our use of IGI asks how intelligence relates to people: whose purposes it serves, what it remembers, how authority is assigned and how decisions can be challenged. It is not a scientific ranking after ASI.
A conceptual discussion map, not an inevitable timeline.
IGIQANT’s proposed design orientation: intelligence that grows with humanity. Apply the question of purpose, accountability and control at every capability level.
IGI is not an established scientific category or a demonstrated stage beyond ASI.
For a first prototype, that orientation would mean an explicit human purpose, governed memory, bounded agent work, external action permissions and review of both outcomes and failures. The architecture would need to be implemented and tested. None of these components makes the others unnecessary.
Collaboration itself is not a safety result. A 2024 human–AI meta-analysis found that, on average across the studied tasks, combinations performed worse than the better of humans or AI alone. Results varied. For IGIQANT, the implication is a testable requirement: show when a particular collaboration improves outcomes, and reject arrangements that merely add confidence or workload.
Nor is one user’s satisfaction the same as alignment with humanity. A system can assist its immediate user while harming another person. Our design questions therefore include consent, distribution of benefits, appeal, privacy and the limits that remain binding even when a user asks otherwise.
The answer to Coxon’s warning cannot be “trust our vision.” It must be a willingness to expose that vision to serious tests, publish limitations and stop when the evidence does not support further autonomy. Development with humanity includes the human authority to decide that a particular form of development should not proceed.
Read the proposed IGI design principles Explore the research questionsSources & editorial method
- The original warningJacob Coxon · X thread
Primary source: opening post and six follow-up posts read in the browser on 10 September 2026. His forecasts and reports of others’ beliefs are attributed, not treated as established facts.
- Coxon’s resignationArs Technica · 9 September 2026
Secondary reporting used to cross-check the news context; technical claims rely on primary sources.
- Pacing the FrontierAdvocacy statement · July 2026
Primary advocacy statement. Supports attribution of its position, not proof that its forecasts will occur.
- AI Safety Report 2026International AI Safety Report · Section 2.2.2
Scientific synthesis published February 2026; a dated assessment, not a live verdict on September systems.
- AI ControlGreenblatt et al.
Research on safety protocols under intentional subversion in a specified programming-task setting.
- AI Risk Management FrameworkNIST · AI RMF Core
Voluntary risk-management framework; no claim of IGIQANT certification.
- Excessive agencyOWASP · LLM06:2025
Security guidance on functionality, permissions and autonomy.
- The AI ActEuropean Commission
Official overview of a risk-based legal framework. Specific obligations require assessment of the actual system and role.
- Levels of AGIMorris et al.
Research framework distinguishing performance, generality and autonomy; not a universal classification.
- Human–AI collaborationVaccaro, Almaatouq & Malone · 2024 meta-analysis
2024 systematic review and meta-analysis; results depend on the tasks and systems studied.
Prepared with AI assistance for IGIQANT. Original source review: 10 September 2026. Presentation and selected references reviewed on 15 September 2026. Research findings, attributed positions and IGIQANT proposals are distinguished throughout. No endorsement by cited researchers or institutions is implied. Corrections and substantive counterarguments are welcome at contact@igiqant.com.