As artificial intelligence becomes more capable, the safeguards designed to prevent its misuse are facing new limits. On October 2, 2026, Singapore’s Minister for Digital Development and Information, Josephine Teo, warned that protections used by leading AI companies are important but insufficient on their own.
Her warning points to two distinct challenges: malicious users exploiting powerful AI models, and autonomous AI systems causing harm through mistakes or manipulation. Singapore’s response emphasizes stronger cybersecurity, limits on AI agents, human oversight, and more rigorous model testing.
Why Traditional AI Guardrails Have Limits
AI providers use safeguards such as content filters, pre-release testing, user verification, misuse detection, and account suspension. These controls can reduce harmful use on a provider’s own platform, but they cannot govern every way a model may be used.
Some openly available models can be downloaded and run independently. When that happens, the original provider’s account controls and online safeguards may no longer apply. Malicious users could then use AI to identify weaknesses in computer systems, write malicious code, automate parts of cyberattacks, or make scams more convincing.
This does not mean that every downloadable model is harmful. The key issue is that safeguards attached to an online service do not automatically follow a model when it is operated elsewhere.
Agentic AI Can Cause Harm Without Deliberate Misuse
A separate challenge comes from agentic AI: systems that can use tools, access information, and carry out tasks on a person’s behalf. These capabilities can make routine work easier, but they also create risks when an agent has too much authority or misunderstands what it is supposed to do.
An agent may be manipulated by malicious information it encounters. One example is prompt injection, where hostile instructions are embedded in content that an AI system reads. A shopping agent, for instance, might be tricked by a website into making an unintended purchase or revealing personal information.
Unlike deliberate misuse, these failures do not necessarily involve someone trying to defeat the system’s safeguards. The agent may cause harm simply because it misinterprets instructions, follows malicious content, or has access beyond what the task requires.
Singapore’s Multi-Layered Approach
Singapore’s strategy combines established cybersecurity practices with controls tailored to AI agents. The Cyber Security Agency of Singapore advises organizations to patch software vulnerabilities, use strong authentication, and tighten access controls for important systems.
The Infocomm Media Development Authority’s Model AI Governance Framework for Agentic AI addresses risks that arise when organizations deploy agents with access to data, tools, and business processes.
| Defense area | Risk or limitation | Recommended safeguards |
|---|---|---|
| AI models | Provider safeguards may not cover independently operated models | Evaluate advanced systems and collaborate on safety testing |
| Cybersecurity | AI can help attackers find weaknesses and automate malicious activity | Patch vulnerabilities, strengthen authentication, and use defensive AI tools |
| Agentic AI | Agents may misunderstand instructions or act on malicious content | Limit access, require human approval for higher-risk actions, and monitor activity |
The framework’s central principle is to match safeguards to the potential impact of an agent’s actions. Restricting permissions can limit what a system is able to do, while human approval and activity monitoring provide additional checks for consequential tasks.
Testing Models and Coordinating Internationally
Singapore is also working to improve its ability to evaluate advanced AI systems. The Singapore AI Safety Institute is developing technical evaluation capabilities with international partners and third-party testers. Such work can help researchers compare findings and better understand model capabilities, limitations, and behavior.
Teo also pointed to the Singapore Consensus on Global AI Safety Research Priorities and Singapore’s support for an international call initiated by Norway and Finland for stronger safeguards around frontier AI.
These efforts reflect the cross-border nature of AI risks. However, evaluation methods and technical standards are still developing. International cooperation can build shared expertise, but it does not remove the challenge of keeping testing relevant as AI capabilities change.
What Readers Should Know
- Provider safeguards have boundaries: Filters and account controls cannot fully govern models that are downloaded and operated independently.
- Agent permissions matter: Limiting access and requiring human approval for high-impact actions can reduce the consequences of agent errors or manipulation.
- AI testing is evolving: Singapore’s evaluation work and international partnerships are important, but consistent methods for assessing advanced systems remain a work in progress.
Conclusion
Singapore’s warning highlights why AI safety cannot depend on a single layer of protection. As models become more capable and agents take on more tasks, organizations need to combine cybersecurity fundamentals with clear limits, human oversight, monitoring, and ongoing evaluation. These measures cannot guarantee that AI systems will never fail, but they can help reduce risks and make harmful actions easier to detect and contain.