900 vulnerability hypotheses and 12 billion tokens.
"Just write the right prompt," they said.
The project by the numbers:
๐ 900 vulnerability hypotheses proposed by models.
๐งช 69 findings confirmed with working PoCs.
๐ 103 groups of findings submitted to developers; fixes already accepted for 46.
๐ค 1,173 human prompts and 615 sub-agents.
⏱️ Approximately 210 agent-hours of work.
๐ฅ 12.16 billion tokens processed.