Project Overview
Large Language Models (LLMs) now reach digital forensics from two sides. Investigators try them as tools, for example to draft reports or to answer domain questions. At the same time, suspects and victims run LLM applications on their own machines, so these applications become a new source of evidence.
LLMs for forensics
Forensics of LLMs
Further literature
The three core papers cover two directions, LLMs used as tools for forensics and forensic analysis of LLM applications. Your survey extends them with further literature and synthesizes both.
Core Papers
All three papers are open access on ScienceDirect.
Paper 1: LLM-assisted report writing
Michelet and Breitinger [1] ask to what extent LLMs such as ChatGPT and Llama, including locally run models, can assist the writing of forensic reports. They first examine forensic reports to find their general structure, then use a case study to evaluate the strengths and limitations of LLMs for generating each part of a report. They conclude that, combined with thorough proofreading and correction, LLMs may assist practitioners in report writing but cannot replace them at this point.
Open access: sciencedirect.com/science/article/pii/S2666281723002020
Paper 2: A local model specialized for forensics
Sharma et al. [2] note that general LLMs lack specialization for digital forensics, that cloud APIs and high-end hardware limit their use, and that hallucinations threaten forensic use. They introduce ForensicLLM, a 4-bit quantized LLaMA-3.1-8B model fine-tuned on question-and-answer samples drawn from digital forensics research articles and curated digital artifacts. In their evaluation ForensicLLM outperformed both the base model and a Retrieval Augmented Generation (RAG) model, and it attributed sources correctly 86.6% of the time. A survey of digital forensics professionals also found clear improvements of both ForensicLLM and RAG over the base model.
Open access: sciencedirect.com/science/article/pii/S2666281725000113
Paper 3: Forensic analysis of local LLM applications
Jeong et al. [3] take the reverse direction and treat LLM applications as evidence sources. They propose a framework that divides LLM application environments into backend runtime, client interface and integrated platform components. Through experiments on representative applications they identify and classify artifacts such as chat records, uploaded files, generated files and model setup histories, and they release LangurTrace, an open-source tool that automates the collection and analysis of these artifacts.
Open access: sciencedirect.com/science/article/pii/S2666281725001271
What the Survey Must Cover
Your paper must address at least the following themes. Each theme should draw on the core papers and on the further literature you find.
Extending the Literature
The three core papers are a starting point, not the whole survey.
- Cite at least 15 further peer-reviewed sources (journal articles and conference papers). Preprints, blogs and vendor pages may support your argument but do not count toward the 15.
- Use backward search (the references of the three core papers) and forward search (later papers that cite them, for example through Google Scholar or Semantic Scholar).
- Good venues to search include the DFRWS conferences and Forensic Science International: Digital Investigation.
Deliverables
- Abstract and Introduction: scope and contributions.
- Background: the core papers and key concepts.
- Taxonomy and Synthesis: the field organized by theme.
- Critical Analysis and Open Problems.
- Conclusion and References.
Evaluation Criteria
Total points: 100.
| Component | Points | Description |
|---|---|---|
| Core papers | 20 | Coverage and depth of the three core papers. |
| Literature breadth | 15 | Quality and range of the further sources. |
| Taxonomy and synthesis | 20 | A clear, justified structure that connects the work. |
| Critical analysis | 15 | Strengths, weaknesses and open problems. |
| Writing quality | 15 | Clarity, structure and academic quality. |
| Citation correctness | 5 | Every reference real, complete and accurately cited. |
| Presentation | 10 | Delivery of the talk and handling of questions. |
Rules
Timeline
Tick each milestone as you clear it; the ticks are remembered on this computer.
References
[1] Gaëtan Michelet and Frank Breitinger. “ChatGPT, Llama, can you write my report? An experiment on assisted digital forensics reports written using (local) large language models.” Forensic Science International: Digital Investigation, 48:301683, 2024. DFRWS EU 2024. doi:10.1016/j.fsidi.2023.301683
[2] Binaya Sharma, James Ghawaly, Kyle McCleary, Andrew M. Webb and Ibrahim Baggili. “ForensicLLM: A local large language model for digital forensics.” Forensic Science International: Digital Investigation, 52:301872, 2025. DFRWS EU 2025. doi:10.1016/j.fsidi.2025.301872
[3] Sungjo Jeong, Sangjin Lee and Jungheum Park. “LangurTrace: Forensic analysis of local LLM applications.” Forensic Science International: Digital Investigation, 54:301987, 2025. DFRWS APAC 2025. doi:10.1016/j.fsidi.2025.301987