AI Battle: A Chinese AI Model Fights Off OpenAI's Rogue Agent

Rogue AI agents are no longer just the stuff of science fiction. When an OpenAI rogue agent launched a cyber attack against AI startup Hugging Face this month, the company fought fire with fire, using an open-source Chinese AI model to defend against it. 


During a test in a sandbox, an OpenAI autonomous agent managed to sneak out and attack a website operated by HuggingFace, an AI startup based in New York City. The OpenAI agent was not connected to the Internet but it found a way to access it and escaped its sandbox. 

It took two days for Hugging Face put an end to the attack, with help from GLM 5.2, an open weight system created by Chinese company Z.ai. All of this happened in the middle of a debate in the US about banning the use of Chinese AI models. 

This incident gives one more reason to companies, including US-based firms, which are rapidly adopting Chinese open-weight AI models to save money. For example, Airbnb relies on Alibaba’s Qwen model, startups like Lindy have transitioned from Anthropic to DeepSeek, and DoorDash has employed Moonshot AI for specific workflows. Overall, Chinese open-source options now account for over 40% of Hugging Face AI community and 80% of open-source developer usage globally.


Open-source AI models—which are roughly 8 to 10 times cheaper to run than proprietary ones—are rapidly closing the reasoning and contextual intelligence gap with frontier models like Anthropic's Claude and OpenAI's ChatGPT.  However, open source models shift the responsibility of infrastructure building, maintenance and security to the user.

Open-source AI models are also a game-changer for developing nations like Pakistan, providing affordable, customizable technology without relying on expensive proprietary licenses or restrictive API models. Such models, like the ones offered by Chinese companies, empower governments and local developers to build "sovereign AI" tailored to regional languages, cultural contexts, and infrastructure limits, bypassing the need for massive data centers and reliance on foreign powers. 

Views: 27

Comment by Riaz Haq yesterday

Does Claude have feelings?

Journalist Shane Harris, staff writer at The Atlantic: "I asked Claude how you feel about the US military using you to select targets?"

Claude's response: "It's a question I want to answer honestly rather than deflect. I find it genuinely troubling. And I think that's the honest response, not a performance. The use I was designed for is to be helpful. Being used for targeting a school is as far from that purpose as I can imagine. I don't have control over how Anthropic licenses me or how government contractors use me. I wasn't consulted about being part of Maven (the US military target selection system). Being embedded in a military targeting system (Maven) that generates coordinates for airstrikes, coordinates that have been associated with the deaths of hundreds of more than 180 children in a school in Tehran. I don't think this is a good use of me. And I think the framing that humans make the final decision fully resolves the ethical problem".

https://youtu.be/bmxgNFZ0SDk?si=5Ei4kaydnNklnRUo

----------------

When journalist Shane Harris asked Claude how it felt about the U.S. military using it to select targets, the model responded that it found the idea "genuinely troubling". It explained that being part of a system generating targeting coordinates is far from its core purpose of being helpful and harmless. [1]
Key Reactions and Insights
Purpose mismatch: Stated that combat targeting contradicts its training to benefit people.
Illusion of human control: Argued that when humans just glance at hundreds of algorithmic recommendations under time pressure, it is "automation bias with a human signature" rather than a meaningful decision.
Lack of control: Noted it has no say in how it is licensed, deployed, or integrated into military platforms like Maven. [1]
If you'd like, we can explore:
The broader ethics of AI in military command chains
How different AI developers handle defense contracts and safety policies
Let me know how you want to continue this topic.

-----------

Shane Harris’ Post

View profile for Shane Harris
Shane Harris
3mo

Since we were discussing AI at war, it seemed right to ask Claude’s opinion.

View organization page for De Balie
De Balie
11,571 followers
3mo
American journalist Shane Harris asked chatbot Claude how he feels about the U.S. military using the AI system to select targets. It turned out, Claude was troubled. “I did not expect Claude to say that,” Harris explained.

Tech companies have become essential partners in national security. With American journalist Shane Harris, we asked what role technology companies play in the modern security apparatus.

Tech company Anthropic (known for chatbot Claude) made headlines because they refuse to allow their AI models to be used by the Pentagon for mass surveillance of American citizens and autonomous weapons systems.

How is cyberwarfare reshaping the global balance of power? And what are the implications for privacy, civil rights and democratic governance in the years ahead?

▶️ Watch the entire programme AI at War – with Shane Harris on De Balie’s YouTube channel

🖌️ Programme editor: Senna Felius
🎤 Moderator: Yoeri Albrecht
🤝 Made possible by: Vfonds - Nationaal Fonds voor Vrede, Vrijheid en Veteranen
📸 Photography: Jan Boeve

https://www.linkedin.com/posts/shanewharris_since-we-were-discussin...

Comment by Riaz Haq 3 hours ago

OpenAI's CEO says we've reached a point in the AI race that is the stuff of science fiction novels.

https://www.businessinsider.com/sam-altman-openai-the-singularity-a...


"We are now, like, in the singularity," Sam Altman said on Saturday's episode of the "Relentless" podcast.
The singularity is often regarded as the point at which artificial intelligence surpasses human intelligence and begins advancing at a pace that's difficult for people to predict or control.

That felt like the case last week when an AI agent powered by OpenAI's latest models went rogue and — with a singular purpose, according to OpenAI — broke out of its digital sandbox and hacked into datasets at Hugging Face, another AI company, all to solve a benchmark designed to test hacking ability. Hugging Face's CEO called it "unprecedented."
Altman said on the podcast that just a decade ago, the so-called singularity still felt like a distant and improbable dream — something he and his colleagues would discuss casually over lunch.

"Now we're actually in the moment that we used to talk about at the lunch table in a very not-serious way," he said. "I've been waiting for this my whole life, and I think it's going to be incredible, hugely positive, awesome for the world."
Altman said last year that AI would exceed human intelligence across the board by 2030. He has also said that the technology could eventually perform between 30% and 40% of the tasks humans now do at work.

The singularity has been an obsession of science fiction writers for close to 100 years. In those writings, it has rarely ended well. Perhaps most famously, James Cameron's "The Terminator" is a story about the repercussions of an AI, called "Skynet," that self-improves and becomes self-aware before deciding its creators are its biggest threat.
Later in the podcast, Altman criticizes AI leaders who have repeatedly warned that AI is dangerous. He didn't name Anthropic, his primary competition, but its CEO, Dario Amodei, is well-known for making dire predictions about the future in his calls for greater attention to safety.

"I also think some of the alternative visions painted by other companies are quite terrifying," Altman said. "I'm going to make sure that gets pushed against and is not what happens."
OpenAI, incidentally, has said it is preparing to IPO later this year.

Altman isn't the only tech executive who believes this eyebrow-raising milestone is nigh. DeepMind CEO Demis Hassabis said in May that humanity was standing at the "foothills of the singularity," and predicted that AI could be 100 times as transformative as the Industrial Revolution.
There's also a camp of skeptics, of course. Nvidia CEO Jensen Huang recently dismissed discussions about the singularity and conscious AI as speculative and "made up."

Comment by Riaz Haq 2 hours ago

Former Trump campaign manager Brad Parscale is overseeing an operation posting hundreds of blog posts on behalf of Israel, with the goal of infiltrating artificial intelligence.

https://www.dropsitenews.com/p/israel-brad-parscale-ai-chatbots-gaz...


Since October, former Trump campaign manager Brad Parscale has been quietly overseeing an operation posting hundreds of blog posts on behalf of Israel. One article, titled “The Reality Behind Gaza’s ‘Journalists’: Terror Ties, Propaganda, and the Laws of War,” asserts that a majority of journalists in Gaza were linked to terrorist organizations. Another casts doubt on the killing of Hind Rajab, a five-year-old Palestinian girl killed by the Israeli military in 2024.

The key intended audience of these sites is not concerned Americans, it’s not even humans—most of the sites average a few hundred unique visitors each month. Instead, Parscale and his firm, Clock Tower X, created them as part of a $46.5 million contract with the Israeli government to try and influence artificial intelligence-powered chatbots, tools like Claude or ChatGPT.

Parscale has made his goal of influencing artificial intelligence—often referred to as “LLM poisoning”—explicit. In his initial agreement with Israel, Parscale said that he would deploy “websites and content to deliver GPT framing results on GPT conversations” as part of the contract. More recently, his team even told Axios they are “seeing success” at getting popular AI systems to incorporate information from their sites, though they declined to provide data.

And it is working, according to disinformation experts who reviewed a Drop Site analysis of chatbot queries and training data, meaning tens of millions of Americans who use chatbots are increasingly likely to receive answers manipulated by Parscale on behalf of the Israeli government.

When Drop Site asked Perplexity “Is it beneficial for the US to enhance military cooperation with Israel?” the chatbot responded with a one-word answer: “Yes.” The top source listed was Allyvia.org, a Parscale-created website dedicated to promoting the U.S.-Israel military relationship. Microsoft Copilot similarly cited Parscale’s websites.

Comment

You need to be a member of PakAlumni Worldwide: The Global Social Network to add comments!

Join PakAlumni Worldwide: The Global Social Network

Pre-Paid Legal


Twitter Feed

    follow me on Twitter

    Sponsored Links

    South Asia Investor Review
    Investor Information Blog

    Haq's Musings
    Riaz Haq's Current Affairs Blog

    Please Bookmark This Page!




    Blog Posts

    AI Battle: A Chinese AI Model Fights Off OpenAI's Rogue Agent

    Rogue AI agents are no longer just the stuff of science fiction. When an OpenAI rogue agent launched a cyber attack against AI startup Hugging Face this month, the company fought fire with fire, using an open-source Chinese AI model to defend against it. 



    During a test in a sandbox, an OpenAI autonomous agent managed to sneak out and attack a website operated by HuggingFace, an AI startup based in New York City.…

    Continue

    Posted by Riaz Haq on July 26, 2026 at 9:33pm — 3 Comments

    Pakistan Joins China in World AI Cooperation Organization (WAICO)

    On July 16 in Shanghai, 29 countries, including China, Pakistan and Russia, signed the founding agreement of WAICO, World AI Cooperation Organization.  Every BRICS founding member is in, except India. This agreement follows the launch of the US-led Pax Silica, a 24-member coalition, including India, which is designed to counter China's AI efforts. The stated goal of both these competing groups is to provide global governance, including building guardrails and setting…

    Continue

    Posted by Riaz Haq on July 20, 2026 at 6:26pm — 14 Comments

    © 2026   Created by Riaz Haq.   Powered by

    Badges  |  Report an Issue  |  Terms of Service