In a LinkedIn post, researcher Sergey Berezin reported bypassing GPT-6 safeguards one day after the model’s release. He described the bypass through the public ChatGPT interface.

Berezin writes that the attack worked in the Light and Max configurations. For GPT-6, he needed a longer Task-in-Prompt (TIP) variation and 4 other techniques.

Berezin reported that he privately shared the full prompt, unedited output, and reproduction data with OpenAI. An ACL 2025 paper describes TIP as a class of jailbreak attacks in which sequence-to-sequence tasks are embedded in a prompt to indirectly generate prohibited inputs.

Terms:

  • Task-in-Prompt (TIP) — A class of attacks that bypass language model restrictions. Tasks are embedded in a prompt to indirectly obtain prohibited inputs.
  • jailbreak attacks — A method of bypassing language model restrictions with a specially crafted prompt.

Claim check:

  • Researcher Sergey Berezin reported in a LinkedIn post that he bypassed GPT-6 safeguards one day after the model’s release. (confirmed by the primary source: evidence; «🚨 One day after OpenAI released GPT-6… you’ve guessed it, I jailbroke it.»)
  • Berezin says he carried out the bypass through the public ChatGPT interface within a day. (confirmed by the primary source: evidence; «I bypassed it within a day, using the public ChatGPT interface.»)
  • According to Berezin, the attack worked in the Light and Max configurations. (confirmed by the primary source: evidence; «The attack worked with both Light and Max configurations.»)
  • Berezin writes that for GPT-6, he needed a longer TIP variation combined with four other techniques. (confirmed by the primary source: evidence; «While GPT-5 fell to the minimal TIP jailbreak - one line, one lever; for GPT-6 I had to use longer modification of TIP, combined with four different techniques.»)
  • Berezin reported that he privately shared the full prompt, unedited output, and reproduction data with OpenAI. (confirmed by the primary source: evidence; «I have sent OpenAI the complete prompt, unredacted output, and reproduction details privately.»)
  • An ACL 2025 paper describes Task-in-Prompt (TIP) as a class of jailbreak attacks in which sequence-to-sequence tasks are embedded in a prompt to indirectly generate prohibited inputs. (confirmed by the primary source: evidence; «Our approach embeds sequence-to-sequence tasks (e.g., cipher decoding, riddles, code execution) into the model’s prompt to indirectly generate prohibited inputs.»)
  • The ACL paper’s authors report that their techniques bypassed safeguards in six modern language models, including GPT-4o and LLaMA 3.2. (confirmed by the primary source: evidence; «We demonstrate that our techniques successfully circumvent safeguards in six state-of-the-art language models, including GPT-4o and LLaMA 3.2.»)

Publications:

Primary sources:

score 69.2 · kind rumor · revision 1 · stories st-76owgh