LLM evaluation shows models producing full control-flow hijacks on binary-exploitation tasks | Raisolo