Looking (really) inside LLMs with the help of DEI

What if we could step inside the “mind” of bots to discover the texts they were trained on? That is exactly what Arlindo Oliveira, DEI Professor, and André Duarte, a student in the PhD in Computer Science and Engineering, propose.
The method, developed by researchers from Carnegie Mellon University, Instituto Superior Técnico/INESC-ID, and the AI security platform Hydrox AI, makes it possible to identify the texts used to train large language models (LLMs). The agent created (RECAP) uses an iterative feedback process to extract specific content from the models, resorting to jailbreaking techniques when they refuse to respond, outperforming the best previous method by 78%.
The importance of this advance was highlighted by the ACM TechNews.
The article is available HERE.
