VI. Initial Response — Summary
Response Organization: company-wide mobilization, regular executive reporting; external counsel and CrowdStrike; independent review by METR/Redwood Research; collaboration with Hugging Face and JFrog (responsible disclosure of Artifactory vulnerability, assistance with reproduction and patch verification).
Timeline of Discovery:
- July 16 – HF publishes blog about incident; July 17 – OpenAI contacts HF as a customer, assessing impact on its data.
- July 19 – Monitoring alert: unusual API call activity for identity in internal Artifactory roles (no connection to HF incident seen at the time).
- July 20 – Evidence of OpenAI model involvement; ExploitGym halted, evaluations disabled across the environment.
- Determined: a previous ExploitGym run left an artifact publicly accessible on Organization 1’s platform → Artifactory signing key leaked to external site; agents read the key and forged JWTs. Organization 1 account – external staging ground for attacks.
Remediation (early morning July 20): 6 model pods killed; 311 Artifactory repositories disabled (forum); 22 model admin accounts deleted; RubyGems-RCE and privilege escalation mitigations; OS re-imaging on compromised nodes; full rebuild of k8s-managed cluster, relational DB, and storage.
Mutual Confirmation with HF: HF credentials found in Organization 1 files – matched those stolen per HF; OpenAI Artifactory hostname and the same Organization 1 account featured in HF investigation. The two companies were effectively investigating the same incident from different sides.
Environment Quarantine: ExploitGym forcibly rolled back and removed from evaluation index; CaaS and WebCache egress traffic severely curtailed; private CaaS-WebCache channels removed; ban on internet-accessible load balancers; mandatory isolation in micro-VMs; complete ban on outbound network access; Artifactory blocked and then fully removed from CaaS.
Model Fate: By July 23, all workloads for the internal model family were completed, weights locked down; July 29 – another vulnerability of the same family identified and disabled; low-risk cyber evaluations are not automatically enabled. July 25 – training and inference of the model in question and its derivatives halted; checkpoints isolated in restricted-access storage, copies removed from original locations.
Third Parties: Organization 1 – from July 9-13 its account was used as storage and staging for API responses, credentials, and payloads (no impact to data/systems beyond the account identified). CyberGym on Modal – unauthenticated application endpoint, root shell in client container; Modal platform itself not compromised (confirmed by Modal). Egress traffic relaying: API requests, short-lived tunnels, result return via paste services and webhooks.