Researchers from IIT Bombay and Adobe Research have developed a method called Previous-Token Prediction (PTP) that can reconstruct a model's original prompt from its generated text alone. This technique trains an inverse model to predict preceding tokens instead of the next one. A PTP model trained on a small open-source model like Qwen-3-0.6B can accurately infer the intent and meaning of prompts sent to larger, proprietary models such as GPT-4o, even without knowing which model produced the output. While the current demonstration is limited to short prompts, it raises concerns about the security of proprietary system prompts used by companies. AI
IMPACT Raises concerns about the security of proprietary system prompts and could impact how companies protect their AI instructions.
RANK_REASON Research paper detailing a new method for prompt reconstruction. [lever_c_demoted from research: ic=1 ai=1.0]
- Adobe Research
- Anthropic
- Claude
- GPT-4o
- Indian Institute of Technology Bombay
- Indian Institutes of Technology
- Previous-Token Prediction
- Suhail
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →