Fix RCE for canary exploit
Fixed in: public source commit; a separate fixed release is not established
details
- Finding IDs
- F-PURPLELLAMA-003
- Status
- fixed on main (release not confirmed)
- Fixed in
- public source commit; a separate fixed release is not established
- Reported via
- security contact
- Note
- The canary-exploit benchmark parsed model-response values with eval(), allowing input data to be evaluated as Python when that benchmark path ran. Commit 48fa920b of 11 March 2026 replaces those evaluations with ast.literal_eval() in verify_response.py. This confirms the recorded code/data-boundary change in public source, not a separately identified stable release or deployment. The existing finding is retained once, without adding another result for this recheck. A maintainer confirmation tying the patch to this report and public reporter credit remain unconfirmed.
F-PURPLELLAMA-003: Code / markup injection. Untrusted values reach eval instead of literal-only parsing. Reviewed 24 Sep 2026. Mechanism assessed by 1seal.
Mechanism source for F-PURPLELLAMA-003
Security area (1seal assessment): Semantics. Replacing eval with literal parsing keeps model-response data from being interpreted as executable Python. Reviewed 26 Sep 2026.