OpenAI Pauses Internal Work on Astra Model Over Fears It Could Launch Cyberattacks Autonomously
OpenAI has halted some internal activities involving its new Astra model after concluding it cannot yet rule out that the system has reached a 'Critical' cybersecurity risk threshold, intensifying scrutiny of frontier AI safety just as ChatGPT and Google's Gemini both pass one billion users.
Data centre server racks with abstract AI network visualisation overlay
What happened?
OpenAI has paused some internal activities involving its newest frontier model, code-named Astra, after determining that the company cannot yet rule out that the system has reached what it internally classifies as a 'Critical' cybersecurity risk threshold — meaning the model could potentially be capable of autonomously identifying and executing cyberattacks, according to reporting from CNBC.
The disclosure comes in a week of intensifying scrutiny of how frontier AI companies handle both safety and data practices. The BBC reported that Twitch, the Amazon-owned livestreaming platform, has faced backlash after enabling an opt-out (rather than opt-in) system allowing Amazon to use streamers' content to train AI models, while a separate report from Eastern Herald described a new OpenAI feature called 'Computer History' that logs users' keystrokes, clicks and application switches to build long-term memory for ChatGPT, raising fresh privacy concerns.
Separately, The Verge reported that both ChatGPT and Google's Gemini have now surpassed one billion users each, underlining how rapidly frontier AI products have scaled into mainstream daily use even as the underlying safety and privacy questions remain unresolved.
Key points
- OpenAI paused some internal activities involving its Astra model over concerns about autonomous cyberattack capability.
- The company says it cannot yet rule out the model has reached its internally defined 'Critical' cybersecurity threshold.
- Twitch/Amazon face backlash over an opt-out system letting Amazon train AI on streamers' content, per the BBC.
- OpenAI's new 'Computer History' feature logs keystrokes and app activity to build ChatGPT memory, raising privacy concerns.
- ChatGPT and Google Gemini have both passed one billion users each, per The Verge, intensifying the stakes of safety and privacy decisions at scale.
- U.S. and EU regulators are increasing oversight efforts around frontier-model security evaluations.
What we know
According to CNBC, OpenAI uses an internal risk-tiering framework to evaluate new models before and during deployment, with 'Critical' representing the highest tier of concern for categories including cybersecurity, biological and chemical risk, and autonomous replication. The company's decision to pause certain internal activities involving Astra — reportedly including some testing or development workflows — reflects an unusually direct and public acknowledgement that a model in active development may have crossed, or come close to crossing, that threshold specifically in its potential capacity to autonomously plan or carry out cyberattacks.
OpenAI has not disclosed the specific technical evidence behind the assessment, and the company's framing — that it 'cannot yet rule out' Astra reaching the Critical threshold — is itself notable for its caution: it is not a confirmed finding of dangerous capability, but a statement of uncertainty serious enough to justify halting some work while further evaluation takes place. CNBC notes this sits alongside a broader pattern of AI evaluation incidents and growing U.S. and EU regulatory oversight of frontier model security.
On the separate privacy front, the Eastern Herald report describes OpenAI's 'Computer History' feature as a way of building richer, more persistent memory for ChatGPT by observing a user's ongoing computer activity, rather than relying solely on typed conversations. The report raises concerns about the sensitivity of data such a feature could capture, including banking and medical-related activity conducted on a device where the feature is active.
Background
OpenAI, along with rivals including Google DeepMind and Anthropic, has for several years published structured frameworks — often called 'preparedness' or 'responsible scaling' policies — that define risk tiers across categories such as cybersecurity, biological weapons assistance, and models' capacity for autonomous self-improvement or replication. These frameworks were developed partly in response to warnings from AI safety researchers and partly to pre-empt regulatory intervention, by demonstrating that companies are proactively evaluating and, where necessary, restricting their own most advanced systems.
Concerns about AI systems' potential cyberattack capabilities have grown steadily as models have become more proficient at coding tasks, including identifying software vulnerabilities and writing functional exploit code — skills that are dual-use by nature, useful both for defensive security research and for offensive attacks. Independent security researchers and government cybersecurity agencies in the U.S. and Europe have flagged this dual-use dynamic repeatedly over the past two years as one of the more immediate, tangible AI risk categories, in contrast to more speculative long-term risks.
The privacy-related stories this week — Twitch's opt-out AI training policy and OpenAI's keystroke-logging memory feature — reflect a parallel and increasingly prominent front in the AI debate: as companies compete to build more capable and more personalised AI systems, they are increasingly relying on default settings and expansive data collection practices that critics argue place the burden on users to actively opt out, rather than requiring clear, affirmative consent.
Detailed analysis
OpenAI's disclosure regarding Astra is significant less for what it confirms than for what it represents procedurally: a leading AI lab publicly acknowledging that one of its own frontier models may have reached the highest internal risk tier for cybersecurity, and responding by constraining its own internal work rather than only its external deployment. That is a meaningfully different posture from earlier industry practice, in which safety concerns were sometimes only surfaced after external researchers identified concerning behaviour in already-deployed systems.
At the same time, the vagueness of the disclosure — no detail on what specifically Astra demonstrated, what internal activities were paused, or what remediation steps are planned — leaves outside observers with limited ability to independently assess how serious the underlying risk actually is. This tension between transparency and operational security is likely to remain a recurring feature of frontier AI safety disclosures: labs have genuine reasons not to publish details that could itself function as a roadmap for malicious actors, but that same opacity limits independent verification of their safety claims.
The Twitch and 'Computer History' stories, while less dramatic than an autonomous cyberattack risk, point to a related but distinct problem: as AI systems become embedded in everyday digital life at the scale implied by ChatGPT and Gemini's combined two billion-plus users, the sheer volume of personal and behavioural data being captured — often through default-on settings — creates privacy exposure that most users are unlikely to fully understand or actively manage. Regulatory scrutiny in both the U.S. and EU is increasingly likely to treat these consent and default-setting practices as being just as important as headline-grabbing safety thresholds.
Why it matters
If a frontier AI model genuinely approaches or crosses the capability threshold for autonomously identifying and executing cyberattacks, the implications extend well beyond OpenAI's internal development process: such capability, if it proliferates through open-source models, leaked weights, or less safety-conscious competitors, could meaningfully lower the barrier to sophisticated cyberattacks against critical infrastructure, financial systems, and government networks.
The parallel privacy stories matter because they affect a much larger population immediately: with ChatGPT and Gemini together serving more than two billion users, decisions about default data collection settings have an outsized effect on global digital privacy norms, shaping what future generations of users come to consider a normal trade-off for access to powerful AI tools.
What happens next?
Expect continued regulatory interest from both U.S. agencies and the European Union, both of which have been developing more formal frameworks for evaluating frontier AI model risk, potentially including mandatory disclosure requirements for capability thresholds like the one OpenAI has flagged for Astra. Independent AI safety researchers and red-teaming organisations are likely to press OpenAI for more detail on the specific evidence behind its Critical-threshold concern.
On the privacy side, expect continued public and possibly regulatory pressure on both Amazon/Twitch and OpenAI to shift toward opt-in rather than opt-out data collection defaults, following patterns seen in prior controversies over AI training data practices at other major technology companies.
Insight Media Opinion
There is something genuinely reassuring about a leading AI lab publicly disclosing that it cannot rule out one of its own models crossing a critical safety threshold, and responding by restricting internal work rather than quietly pressing ahead. That is exactly the kind of caution safety advocates have been asking frontier labs to exercise for years, and OpenAI deserves credit for erring on the side of disclosure here, however incomplete the details remain.
That said, incomplete disclosure is still a real limitation. The public, regulators and independent researchers are being asked to trust OpenAI's self-assessment without the ability to meaningfully verify it, and that asymmetry becomes more consequential, not less, as these models edge closer to genuinely dangerous capability levels. We would encourage regulators in both the U.S. and EU to move toward requiring a baseline level of independently verifiable evidence behind claims like this one, rather than relying solely on company self-reporting.
On the privacy front, the pattern across both Twitch and OpenAI this week is hard to view charitably: opt-out defaults for AI training and passive activity logging for AI memory both shift the burden of protecting personal data onto users who are unlikely to notice the setting exists, let alone change it. As AI tools embed themselves further into daily digital life, we believe the industry norm needs to move decisively toward opt-in consent as the default, not the exception.
Related Insight Media stories
- The EU's New AI Rules Are Here: What Changed in August 2026
- The Digital Divide Is No Longer Just About Internet Access
- The Age of Electricity: Why AI and Data Centres Are Changing Power Demand
- The Quantum Clock Is Ticking: Why Cybersecurity Teams Are Racing to Prepare
- Edge AI Chips Are Quietly Redesigning Everyday Devices
Sources & further reading
Every claim above can be traced to the documents below.
Author
Insight Media Editorial Desk — original reporting, explainers, analysis and practical guides, researched against primary documents and credible independent reporting. Developing stories are updated when significant new verified information becomes available.