Digital forensic expertise with machine learning to reduce false positives in file recovery
DOI:
https://doi.org/10.57077/monumenta.v14i14.361Keywords:
Digital Forensics, Data Recovery, Machine Learning, False Positives, Information SecurityAbstract
The increasing reliance on computer systems has amplified the importance of digital forensics as a fundamental tool for obtaining reliable technical evidence. In this context, the recovery of deleted files is a critical step, frequently compromised by limitations of traditional tools, especially due to the high occurrence of false positives, which directly impact the reliability of forensic analyses. This work proposes and evaluates a machine learning-based approach to reduce the false positive rate in the deleted file recovery process. The methodology adopted is experimental in nature and involves the construction of a hybrid dataset, composed of synthetic data generated in controlled deletion and fragmentation scenarios, as well as data inspired by real databases. The data are subjected to preprocessing and extraction of forensic attributes, including entropy, byte frequency, and n-grams, and subsequently used to train supervised models and in a hybrid approach based on ensemble learning. The models are evaluated using metrics such as precision, recall, F1-score, and false positive rate (FPR), with direct comparison to traditional heuristic methods. The results indicate that the use of machine learning significantly reduces the occurrence of false positives, contributing to increased reliability in the forensic recovery process. As a contribution, this work presents a reproducible experimental pipeline that integrates computer forensics techniques and machine learning, offering support for improving forensic practices and advancing research in the field.