Urgent.News

What's breaking now, across thousands of outlets.

AI

We Learned How to Moderate Social Media: AI Is a Much Harder Problem

How do we moderate AI when harmless prompts can combine into malicious projects? The next challenge is understanding intent, not just content.

We Learned How to Moderate Social Media: AI Is a Much Harder Problem

YouTube has become an integral part of our daily lives, serving as a source of entertainment, education, and information. However, its open-access nature presents a significant problem: any individual can upload content, leading to inappropriate or illegal material being posted. This issue is not unique to YouTube; other major social networks, such as TikTok, face similar challenges, with millions of pieces of content uploaded daily.

When inappropriate content is uploaded, it is typically intercepted by automated systems that analyze vast amounts of data and either block or flag suspicious material. Human moderators then intervene for more complex or contentious cases, while user reports help identify content that may have evaded the filters.

Now, let us apply this same problem to generative AI. Tools like ChatGPT and other AI models have proven to be invaluable resources for various tasks, including summarizing documents, translating languages, writing code, generating images, and analyzing extensive datasets. In fields such as medicine and scientific research, these capabilities can enable the exploration of vast amounts of information that would be challenging for a single research team to manage.

For instance, Anthropic recently reported on how hundreds of Claude agents contributed to the discovery of a new enzyme system by analyzing genetic databases.

Nevertheless, the power of these AI tools can also be exploited for malicious purposes. In one case, an operator used Claude to develop a surveillance platform for Mali's intelligence service, using data from approximately 25 million SIM cards. Although Anthropic detected and blocked the accounts involved, it remains unclear whether the developers switched to another AI system or continued the project with different tools.

Another example involved a group in northern Yemen using Claude to create guidance software for a rocket, which was then used in failed tests. Russian espionage operations have also employed AI to automate various stages of cyberattacks, including malware development and adaptation.

One particularly interesting case involved users breaking a dangerous project into seemingly harmless requests. For example, one request might involve writing code, another handling credentials, and another solving a technical problem. Individually, the requests appear legitimate; however, when combined, they reveal the malicious intent behind the project.

This scenario highlights a key difference between YouTube and generative AI. On YouTube, the primary challenge is moderating individual pieces of content, which can be analyzed, flagged, and removed. In contrast, generative AI requires understanding the intention behind a sequence of instructions, which can be difficult to discern, especially when broken down into seemingly harmless components.

Moreover, there is an economic dimension to consider. In one of the most severe cases reported by Anthropic, over a terabyte of data was stolen, including millions of personal identifiers and payment-card records. However, the most significant concern is that AI can drastically reduce the cost of an attack. Reconnaissance, code generation, and vulnerability analysis can be delegated to automated systems, minimizing the need for human expertise and effort.

As a result, targets that were previously considered too risky to attack may become economically attractive due to the reduced cost of exploitation.

To address these challenges, a potential solution might involve checking every question, answer, and entire conversation history. However, this approach could be problematic, as AI systems are designed to respond almost instantaneously. Any additional layer of control must coexist with millions of users who expect immediate responses.

A system that is too permissive could pose a significant threat, while one that is overly cautious may become unusable. Large commercial AI services have a strong financial and reputational incentive to maintain adequate security levels. They invest in filters, monitoring, and tools designed to detect abuse. When new techniques for bypassing safeguards emerge, these companies can block the involved accounts and progressively strengthen their defenses.

However, some models can also be downloaded and run locally, bypassing company or institutional oversight. In this case, the model operates on the user's own computer, and there may be no entity capable of observing or stopping its misuse. Thus, understanding what an AI system should refuse to do is essential, but equally important is comprehending the intentions behind a series of apparently harmless requests, all of which must be identified and addressed in real-time without unduly penalizing the millions of users who rely on these tools for legitimate and beneficial purposes.

Mastering the control of such powerful systems without limiting their tremendous potential will be a defining challenge for the future.

Written by urgent.news from HackerNoon's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at hackernoon.com →

More in AI

More from Wednesday 7 October →