LLM Security Leaderboard Now Ranks Every Modality Where Agent Risk Lives: Text, Image, and Audio

Cisco expanded its LLM Security Leaderboard to cover text, image, and audio attacks, adding 102 new multimodal evaluations and bringing the total to 136 models from major providers. The leaderboard measures prompt injection and jailbreak resilience on base configurations, offering developers clearer risk data for agent deployments.

Cover image for LLM Security Leaderboard Now Ranks Every Modality Where Agent Risk Lives: Text, Image, and Audio

Cisco released an updated security leaderboard that now evaluates large language models against attacks delivered via text, image, and audio. The expanded dataset includes one hundred and two new evaluations, bringing the total number of assessed models to one hundred and thirty six. These models come from major providers such as Anthropic, OpenAI, and Google. The aim is to provide organizations with specific data on how models withstand manipulation before they are integrated into autonomous agents. Tests focus on prompt injection and jailbreak attempts to determine if a model can be forced to ignore safety rules. Evaluations use single-turn interactions for image and audio formats, while text assessments may use multi-turn conversations. Models are tested in their base configuration without added guardrails to ensure a consistent baseline for comparison. A combined score averages performance across all formats a model supports, allowing developers to isolate weaknesses in specific modalities. The scores do not reflect the security of an agent in a live production environment, as outside factors like input filtering are excluded from the baseline tests. Results indicate where a model is weak, but they do not guarantee safety in complex, real-world deployments with varied data streams. Some multimodal capabilities are red-teamed internally by providers, which may differ from Cisco’s external testing methods. The data also does not account for future attack vectors that have not yet been developed or deployed.