Skip to main navigation Skip to search Skip to main content

Many-Shot Regurgitation Prompting

Shashank Sonkar, Naiming Liu, Richard Baraniuk

Research output: Chapter in Book/Report/Conference proceedingConference contribution

Abstract

We introduce Many-Shot Regurgitation (MSR) prompting, a new black-box membership inference attack framework for examining verbatim content reproduction in large language models (LLMs). MSR prompting involves dividing the input text into multiple segments and creating a single prompt that includes a series of faux conversation rounds between a user and a language model to elicit verbatim regurgitation. We apply MSR prompting to diverse text sources, including open educational resources textbooks and Wikipedia articles, which provide high-quality, factual content and are continuously updated over time. For each source, we curate two dataset types: one that LLMs were likely exposed to during training (Dpre) and another consisting of documents published after the models’ training cutoff dates (Dpost). To quantify the occurrence of verbatim matches, we employ the Longest Common Substring algorithm and count the frequency of matches at different length thresholds. We then use statistical measures such as Cliff’s delta, Kolmogorov-Smirnov (KS) distance, and Kruskal-Wallis H test to determine whether the distribution of verbatim matches differs significantly between Dpre and Dpost. Our findings reveal a striking difference in the distribution of verbatim matches between Dpre and Dpost, with the frequency of verbatim reproduction being significantly higher when LLMs (e.g. GPT models and LLaMAs) are prompted with text from datasets they were likely trained on. Our results provide compelling evidence that LLMs are more prone to reproducing verbatim content when the input text is likely sourced from their training data. Code is available here.

Original languageEnglish (US)
Title of host publicationArtificial Intelligence in Education - 26th International Conference, AIED 2025, Proceedings
EditorsAlexandra I. Cristea, Erin Walker, Yu Lu, Olga C. Santos, Seiji Isotani
PublisherSpringer Science and Business Media Deutschland GmbH
Pages203-211
Number of pages9
ISBN (Print)9783031984617
DOIs
StatePublished - 2025
Event26th International Conference on Artificial Intelligence in Education, AIED 2025 - Palermo, Italy
Duration: Jul 22 2025Jul 26 2025

Publication series

NameLecture Notes in Computer Science
Volume15881 LNAI
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference26th International Conference on Artificial Intelligence in Education, AIED 2025
Country/TerritoryItaly
CityPalermo
Period7/22/257/26/25

Keywords

  • Large Language Models
  • Membership Inference Attacks

ASJC Scopus subject areas

  • Theoretical Computer Science
  • General Computer Science

Fingerprint

Dive into the research topics of 'Many-Shot Regurgitation Prompting'. Together they form a unique fingerprint.

Cite this