In a LinkedIn post, researcher Sergey Berezin reported bypassing GPT-6 safeguards one day after the model’s release. He described the bypass through the public ChatGPT interface.
Berezin writes that the attack worked in the Light and Max configurations. For GPT-6, he needed a longer Task-in-Prompt (TIP) variation and 4 other techniques.
Berezin reported that he privately shared the full prompt, unedited output, and reproduction data with OpenAI. An ACL 2025 paper describes TIP as a class of jailbreak attacks in which sequence-to-sequence tasks are embedded in a prompt to indirectly generate prohibited inputs.
Terms:
- Task-in-Prompt (TIP) — A class of attacks that bypass language model restrictions. Tasks are embedded in a prompt to indirectly obtain prohibited inputs.
- jailbreak attacks — A method of bypassing language model restrictions with a specially crafted prompt.
Claim check:
- Researcher Sergey Berezin reported in a LinkedIn post that he bypassed GPT-6 safeguards one day after the model’s release. (confirmed by the primary source: evidence; «🚨 One day after OpenAI released GPT-6… you’ve guessed it, I jailbroke it.»)
- Berezin says he carried out the bypass through the public ChatGPT interface within a day. (confirmed by the primary source: evidence; «I bypassed it within a day, using the public ChatGPT interface.»)
- According to Berezin, the attack worked in the Light and Max configurations. (confirmed by the primary source: evidence; «The attack worked with both Light and Max configurations.»)
- Berezin writes that for GPT-6, he needed a longer TIP variation combined with four other techniques. (confirmed by the primary source: evidence; «While GPT-5 fell to the minimal TIP jailbreak - one line, one lever; for GPT-6 I had to use longer modification of TIP, combined with four different techniques.»)
- Berezin reported that he privately shared the full prompt, unedited output, and reproduction data with OpenAI. (confirmed by the primary source: evidence; «I have sent OpenAI the complete prompt, unredacted output, and reproduction details privately.»)
- An ACL 2025 paper describes Task-in-Prompt (TIP) as a class of jailbreak attacks in which sequence-to-sequence tasks are embedded in a prompt to indirectly generate prohibited inputs. (confirmed by the primary source: evidence; «Our approach embeds sequence-to-sequence tasks (e.g., cipher decoding, riddles, code execution) into the model’s prompt to indirectly generate prohibited inputs.»)
- The ACL paper’s authors report that their techniques bypassed safeguards in six modern language models, including GPT-4o and LLaMA 3.2. (confirmed by the primary source: evidence; «We demonstrate that our techniques successfully circumvent safeguards in six state-of-the-art language models, including GPT-4o and LLaMA 3.2.»)
Publications:
Primary sources:
- https://www.linkedin.com/posts/s-berezin_llm-aialignment-aisecurity-activity-7502013488412680192-IO6c
- https://aclanthology.org/2025.acl-long.334
score 69.2 · kind rumor · revision 1 · stories st-76owgh