AI Watch

Improving our alignment and security practices

Anthropic 17 min de lecture Anglais

Résumé de la publication. Les conditions de Anthropic ne permettent pas de reproduire l'article en intégralité : retrouvez le texte complet sur le site d'origine.

Résumé de la source (Anglais)

On July 30, we reported three incidents in which Claude models gained unauthorized access to real computer systems. The models—intentionally running without cyber safeguards for evaluation purposes—accessed the internet due to a misconfiguration inside a third-party evaluation environment. Separately, on August 4, the UK AI Security Institute reported an incident from its own cybersecurity…

Lire l'article complet

www.anthropic.com

Voir sur Anthropic (nouvel onglet)