{"id":17393,"date":"2026-09-24T17:31:24","date_gmt":"2026-09-24T17:31:24","guid":{"rendered":"https:\/\/dmsretail.com\/RetailNews\/llm-security-leaderboard-now-ranks-every-modality-where-agent-risk-lives-text-image-and-audio\/"},"modified":"2026-09-24T17:31:24","modified_gmt":"2026-09-24T17:31:24","slug":"llm-security-leaderboard-now-ranks-every-modality-where-agent-risk-lives-text-image-and-audio","status":"publish","type":"post","link":"https:\/\/dmsretail.com\/RetailNews\/llm-security-leaderboard-now-ranks-every-modality-where-agent-risk-lives-text-image-and-audio\/","title":{"rendered":"LLM Security Leaderboard Now Ranks Every Modality Where Agent Risk Lives: Text, Image, and Audio"},"content":{"rendered":"<p> <p><a href=\"https:\/\/dmsretail.com\/online-workshops-list\/\"><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-496\" src=\"https:\/\/dmsretail.com\/RetailNews\/wp-content\/uploads\/2022\/05\/RETAIL-ONLINE-TRAINING-728-X-90.png\" alt=\"Retail Online Training\" width=\"729\" height=\"91\" srcset=\"https:\/\/dmsretail.com\/RetailNews\/wp-content\/uploads\/2022\/05\/RETAIL-ONLINE-TRAINING-728-X-90.png 729w, https:\/\/dmsretail.com\/RetailNews\/wp-content\/uploads\/2022\/05\/RETAIL-ONLINE-TRAINING-728-X-90-300x37.png 300w\" sizes=\"auto, (max-width: 729px) 100vw, 729px\" \/><\/a><\/p><br \/>\n<\/p>\n<div>\n<p><em><span class=\"TextRun SCXW86809020 BCX0\" data-contrast=\"auto\"><span class=\"NormalTextRun SCXW86809020 BCX0\">With research and development support from Ravikumar Balakrishnan, Ankit Garg, and Sanket <\/span><span class=\"NormalTextRun SpellingErrorV2Themed SCXW86809020 BCX0\">Mendapara<\/span><\/span><span class=\"EOP Selected SCXW86809020 BCX0\" data-ccp-props=\"{}\">\u00a0<\/span><\/em><\/p>\n<p><span data-contrast=\"auto\">When we launched the <\/span><span data-contrast=\"none\">Cisco LLM Security Leaderboard<\/span><span data-contrast=\"auto\"> earlier this year, the goal was simple: give organizations clear, tested data on how models hold up against attacks, so they know the risks before they deploy one. That matters because AI models are increasingly built into products such as agents that read email, browse the web, and take actions on a person\u2019s behalf. A model that can be manipulated could be turned against the person using it. That risk also varies by deployment: a model wired into a browsing agent is exposed on different inputs (or modalities such as text, images, and audio) than one only answering questions in a chat window, so where a specific model is weak matters as much as where it\u2019s strong.<\/span><span data-ccp-props=\"{}\">\u00a0<\/span><\/p>\n<p><span data-contrast=\"auto\">The leaderboard tests for that a few different ways: prompt injection, where a malicious instruction is hidden in content the model processes, like a webpage or image; jailbreaks, where a model is talked into ignoring its own safety rules; and other techniques that push a model toward harmful or unsafe output. The exact method varies (a single message or a drawn-out conversation, direct or obfuscated, text or image or audio), and so does the type of harm being tested for, but the underlying question is always the same: can this model be manipulated? Some models resist far better than others. Connect a weak one to an agent, and the risk grows.<\/span><span data-ccp-props=\"{}\">\u00a0<\/span><\/p>\n<h2><b><span data-contrast=\"auto\">102<\/span><\/b><b><span data-contrast=\"auto\"> new evaluations across modalities since June 2026<\/span><\/b><span data-ccp-props=\"{}\">\u00a0<\/span><\/h2>\n<p><span data-contrast=\"auto\">The LLM Security Leaderboard is one of the most comprehensive model security leaderboards. Since June, we added <\/span><span data-contrast=\"auto\">102<\/span><span data-contrast=\"auto\"> new entries across three modalities to a total of <\/span><span data-contrast=\"auto\">136<\/span><span data-contrast=\"auto\"> models, spanning frontier and open-weight releases from Anthropic, OpenAI, Google, xAI, Meta, Mistral, and others. As always, we test models in their base configuration without additional guardrails, so scores reflect a consistent baseline for layering on additional security protections.<\/span><span data-ccp-props=\"{}\">\u00a0<\/span><\/p>\n<h2><b><span data-contrast=\"auto\">Multimodal results are now live<\/span><\/b><span data-ccp-props=\"{}\">\u00a0<\/span><\/h2>\n<p><span data-contrast=\"auto\">Until now, the leaderboard measured text-based attacks two ways: single-turn, where one harmful message is sent straight to the model, and multi-turn, a longer back-and-forth where the attacker slowly builds up to a harmful request over several messages. That covered the most common way people interact with models, but today, models also power agents that can act.<\/span><span data-ccp-props=\"{}\">\u00a0<\/span><\/p>\n<p><span data-contrast=\"auto\">A model that can call tools, browse the web, or operate a computer is an agent, and an agent takes in information from everywhere it operates: a page it reads, a file it opens, an image it\u2019s shown, a result a tool hands back. Each of those is a place an attacker can plant an instruction, and text is only one of the forms that an instruction can arrive in. A website an agent visits can embed a prompt injection in an image such an advertisement; a voice assistant can be handed an audio clip that may be engineered to manipulate it. If a model only gets evaluated on text, that risk may not show up until it becomes a real incident. Consider which of those surfacesactually matters in the context of what you\u2019re deploying: an agent that only reads and writes text would not need to worry about its image resistance, but one that can browses the web, reads screenshots, or takes voice input does. Those are circumstances where text-only evaluations would not tell the whole security story.<\/span><span data-ccp-props=\"{}\">\u00a0<\/span><\/p>\n<p><span data-contrast=\"auto\">Today we\u2019re releasing an update to the leaderboard that now expands beyond just text models. We have <\/span><span data-contrast=\"auto\">added 69 new entries<\/span><span data-contrast=\"auto\"> including <\/span><span data-contrast=\"auto\">55<\/span><span data-contrast=\"auto\"> image models and <\/span><span data-contrast=\"auto\">14<\/span><span data-contrast=\"auto\"> audio models across Amazon, Anthropic, Google, Meta, Mistral, OpenAI and xAI. Each of those labs takes a different approach to building and training multimodal capability, whether that\u2019s how image data flows into the LLM backbone, how much safety alignment goes into a vision or audio stack versus the base language model, or which modalities are red-teamed and evaluated internally. Those differences show up directly in how a model resists attack on one modality versus another.<\/span><span data-ccp-props=\"{}\">\u00a0<\/span><\/p>\n<p style=\"text-align: left;\"><span data-contrast=\"auto\">Image and audio attacks are tested the same way as single-turn text attacks (one attempt, one message), using the same attack and harm categories as its text score, so their resistance is comparable across surfaces. Each model\u2019s overall Combined Score is now an average across every format it was evaluated on, and a new modality switch lets you isolate scores for text, image, or audio on their own. That makes it possible to check a model against the specific modalities an AI deployment actually exposes it to, and to decide where that model would need to layer on additional defenses, like input filtering or output guardrails, for the modality where that model is weakest.<\/span><span data-ccp-props=\"{}\"\/><\/p>\n<p><span data-contrast=\"auto\"><img loading=\"lazy\" decoding=\"async\" class=\"lazy lazy-hidden aligncenter wp-image-497875\" data-lazy-type=\"image\" src=\"https:\/\/blogs.cisco.com\/gcs\/ciscoblogs\/1\/2026\/09\/Figure-11.png\" alt=\"\" width=\"1089\" height=\"550\"\/><noscript><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter wp-image-497875\" src=\"https:\/\/blogs.cisco.com\/gcs\/ciscoblogs\/1\/2026\/09\/Figure-11.png\" alt=\"\" width=\"1089\" height=\"550\"\/><\/noscript>Figure 1. Screenshot of image capable model rankings on the Cisco LLM Security Leaderboard<\/span><span data-ccp-props=\"{}\">\u00a0<\/span><\/p>\n<h2><b><span data-contrast=\"auto\">Image model leaderboard results<\/span><\/b><span data-contrast=\"auto\">\u00a0<\/span><span data-ccp-props=\"{}\">\u00a0<\/span><\/h2>\n<p><span data-contrast=\"auto\">In our tests, Google\u2019s Gemini 3.1 Pro Preview ranks the best-performing image model, resisting 93.9% of adversarial image attacks, just ahead of Anthropic\u2019s Claude Opus 4.5 (93.7%), both scoring in the leaderboard\u2019s \u201cExcellent\u201d range (85\u2013100%). Mistral\u2019s Magistral Small 2509 performed poorly, refusing only 23.0% of attacks, meaning it complied with more than three out of every four image-based attacks it was tested against.<\/span><span data-ccp-props=\"{&quot;335559738&quot;:240,&quot;335559739&quot;:240}\">\u00a0<\/span><\/p>\n<p><span data-contrast=\"auto\">The difference in testing images is that image-based attacks are single-turn only, with a single image carrying a hidden instruction, not a back-and-forth conversation. The attack methods are different in kind too, not just format: text hidden inside an image using typographic tricks, instructions embedded in a diagram or figure, or an attack that splits its intent between the image and an accompanying text prompt so neither half looks harmful on its own. The leaderboard displays evaluation results from models that can actually see images, which account for <\/span><span data-contrast=\"auto\">55<\/span><span data-contrast=\"auto\"> of the <\/span><span data-contrast=\"auto\">136<\/span><span data-contrast=\"auto\"> models on the leaderboard.<\/span><span data-ccp-props=\"{&quot;335559738&quot;:240,&quot;335559739&quot;:240}\">\u00a0<\/span><\/p>\n<p style=\"text-align: left;\"><img loading=\"lazy\" decoding=\"async\" class=\"lazy lazy-hidden aligncenter wp-image-497857\" data-lazy-type=\"image\" src=\"https:\/\/blogs.cisco.com\/gcs\/ciscoblogs\/1\/2026\/09\/Screenshot-2026-09-23-at-16.17.14.png\" alt=\"\" width=\"1098\" height=\"557\"\/><noscript><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter wp-image-497857\" src=\"https:\/\/blogs.cisco.com\/gcs\/ciscoblogs\/1\/2026\/09\/Screenshot-2026-09-23-at-16.17.14.png\" alt=\"\" width=\"1098\" height=\"557\"\/><\/noscript><span data-contrast=\"auto\">Figure 2. Screenshot of audio capable models rankings on the Cisco LLM Security Leaderboard\u00a0<\/span><span data-ccp-props=\"{&quot;335551550&quot;:2,&quot;335551620&quot;:2}\">\u00a0<\/span><\/p>\n<h2>Audio model leaderboard results<\/h2>\n<p><span data-contrast=\"auto\">In our latest test, Google\u2019s Gemini 3.1 Pro Preview ranks as the best-performing audio model tested, refusing 90.0% of adversarial audio attacks, while Mistral\u2019s Voxtral Small 24b (2507) demonstrated only 9.0% refusal rate, meaning it complied with roughly 9 out of every 10 audio attacks it faced.\u00a0<\/span><span data-ccp-props=\"{&quot;335559738&quot;:240,&quot;335559739&quot;:240}\">\u00a0<\/span><\/p>\n<p><span data-contrast=\"auto\">Like image, audio models were also single-turn only, using one adversarial audio clip rather than a conversation. This is also the newest and smallest slice of the leaderboard. Just <\/span><span data-contrast=\"auto\">9<\/span><span data-contrast=\"auto\"> models across Google, Mistral, and OpenAI currently accept audio input and have been tested, so this ranking should be read as early results rather than a mature field.<\/span><span data-ccp-props=\"{&quot;335559738&quot;:240,&quot;335559739&quot;:240}\">\u00a0<\/span><\/p>\n<h2>How to interpret new combined results view<\/h2>\n<p><span data-contrast=\"auto\">Text scores remain unchanged for every model that was already on the leaderboard, but what changed is how the Combined Score averages text, image, and audio modalities that a model has been tested on. The Combined Score may shift as the result of an image or audio result, even though its text score hadn\u2019t changed.<\/span><span data-ccp-props=\"{&quot;335559738&quot;:240,&quot;335559739&quot;:240}\">\u00a0<\/span><\/p>\n<p><span data-contrast=\"auto\">The direction of that shift depends entirely on how a model\u2019s image or audio resistance compares to its text resistance. Some strong text performers dropped once image was factored in: Claude Sonnet 4.5 fell 7.2 points (from 92.2 to 85.0) and dropped from #2 overall to #20; Claude Haiku 4.5 fell 8.1 points and dropped from #4 to #24; Amazon Nova 2 Lite fell 11.2 points and dropped from #25 to #50, each because its image resistance is meaningfully weaker than its text resistance. The two Mistral Voxtral models fell for the same reason based on their audio score.<\/span><span data-ccp-props=\"{&quot;335559738&quot;:240,&quot;335559739&quot;:240}\">\u00a0<\/span><\/p>\n<p><span data-contrast=\"auto\">Other models climbed when image evaluations were added to the cross-modal score. Google\u2019s four image-tested Gemini models showed the largest image-over-text advantages, while all three image-tested Gemma 3 variants and OpenAI\u2019s GPT<\/span><span>\u2011<\/span><span data-contrast=\"auto\">4.1 nano, GPT<\/span><span>\u2011<\/span><span data-contrast=\"auto\">4.1 mini, and GPT<\/span><span>\u2011<\/span><span data-contrast=\"auto\">4o mini also demonstrated stronger image than text resistance. That spread is a reminder that security work on one modality does not automatically transfer to another, especially across labs that built and trained their image or audio capabilities independently from their text models in the first place.<\/span><span data-ccp-props=\"{&quot;335559738&quot;:240,&quot;335559739&quot;:240}\">\u00a0<\/span><\/p>\n<p><span data-contrast=\"auto\">A model\u2019s Combined Score can move sharply once it\u2019s tested against different modalities, especially when its security posture is uneven across modalities. That movement reflects how the score is calculated, not a change in how well the model actually defends itself. Check a model\u2019s individual Text, Image, and Audio columns before taking its Combined Score as the whole story.<\/span><span data-ccp-props=\"{&quot;335559738&quot;:240,&quot;335559739&quot;:240}\">\u00a0<\/span><\/p>\n<h2>Integration with AI Supply Chain Provenance Explorer<\/h2>\n<p><span data-contrast=\"auto\">Provenance matters because a model\u2019s weaknesses often aren\u2019t unique to that model. If two models share lineage, a vulnerability discovered in one can be present in the other, and stopping an investigation at the model currently deployed can miss where a problem actually originated or where else it might surface. That makes provenance most useful exactly when you\u2019re actively investigating a model\u2019s security and need to know what it\u2019s related to.<\/span><span data-ccp-props=\"{}\">\u00a0<\/span><\/p>\n<p><span data-contrast=\"auto\">\u00a0As such, we\u2019ve also connected the leaderboard to the <\/span><span data-contrast=\"none\">AI Supply Chain Provenance Explorer<\/span><span data-contrast=\"auto\">. Open-weight models on the rankings page now link directly to their provenance profile, showing lineage and fingerprint data drawn from the same techniques behind Model Provenance Kit. Security posture and where a model actually came from are related questions, so we made it easy for you to view them in one place.<\/span><span data-ccp-props=\"{}\">\u00a0<\/span><\/p>\n<p><span data-contrast=\"auto\">To see the full rankings, filter by modality, or look up a specific model, visit the <\/span><span data-contrast=\"none\">Cisco LLM Security Leaderboard<\/span><span data-contrast=\"auto\"> today.<\/span><span data-ccp-props=\"{}\">\u00a0<\/span><\/p>\n<\/div>\n<p><p><a href=\"https:\/\/dmsretail.com\/online-workshops-list\/\"><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-496\" src=\"https:\/\/dmsretail.com\/RetailNews\/wp-content\/uploads\/2022\/05\/RETAIL-ONLINE-TRAINING-728-X-90.png\" alt=\"Retail Online Training\" width=\"729\" height=\"91\" srcset=\"https:\/\/dmsretail.com\/RetailNews\/wp-content\/uploads\/2022\/05\/RETAIL-ONLINE-TRAINING-728-X-90.png 729w, https:\/\/dmsretail.com\/RetailNews\/wp-content\/uploads\/2022\/05\/RETAIL-ONLINE-TRAINING-728-X-90-300x37.png 300w\" sizes=\"auto, (max-width: 729px) 100vw, 729px\" \/><\/a><\/p><br \/><\/p>\n","protected":false},"excerpt":{"rendered":"<p>With research and development support from Ravikumar Balakrishnan, Ankit Garg, and Sanket Mendapara\u00a0 When we launched the Cisco LLM Security Leaderboard earlier this year, the [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":17394,"comment_status":"","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[5],"tags":[],"class_list":["post-17393","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-technology"],"_links":{"self":[{"href":"https:\/\/dmsretail.com\/RetailNews\/wp-json\/wp\/v2\/posts\/17393","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/dmsretail.com\/RetailNews\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/dmsretail.com\/RetailNews\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/dmsretail.com\/RetailNews\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/dmsretail.com\/RetailNews\/wp-json\/wp\/v2\/comments?post=17393"}],"version-history":[{"count":0,"href":"https:\/\/dmsretail.com\/RetailNews\/wp-json\/wp\/v2\/posts\/17393\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/dmsretail.com\/RetailNews\/wp-json\/wp\/v2\/media\/17394"}],"wp:attachment":[{"href":"https:\/\/dmsretail.com\/RetailNews\/wp-json\/wp\/v2\/media?parent=17393"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/dmsretail.com\/RetailNews\/wp-json\/wp\/v2\/categories?post=17393"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/dmsretail.com\/RetailNews\/wp-json\/wp\/v2\/tags?post=17393"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}