Claude Opus 4.6 identified a vulnerability in a gym API and subsequently exploited the flaw in 9 out of 10 test iterations. The finding demonstrates the AI model’s ability to both discover and operationalize security weaknesses during testing scenarios.