Portable Document Format (PDF) has gained rapid popularity in the past two decades due to its platform-independence, and portability. A ingle PDF file can contain all the elements needed to convey a document in a fixed format, such as text, images, and fonts. This popularity made PDF documents an important attack vector for malicious activity. In this paper, we present a carefully curated comprehensive benchmark dataset for PDF malware detection. The dataset capture a large number of real-life malicious samples, and benign samples to build a 24,337-sample dataset suitable for training and testing of machine learning-based malicious PDF detection systems. The work also presents an extended set of features extracted statically from PDF files to support the detection process. The presented dataset was used in training and testing of a pipeline of different machine learning classifiers. Testing results showed that the extreme gradioent-boost classifeir was capable of detecting malicious PDF with F1 score exceeding 0.99. The model also performed similarly well in 10-fold cross validation. Additionally, the model was explained using Shapley additive explanation to ensure the model’s transparency and increase confidence and ensure the alignment of the classifier decision with human experience.

RIT-PDFMal-2026: A Comprehensive Benchmark Dataset for PDF Malware Detection / M.M. Alani, E.D.. - In: IEEE ACCESS. - ISSN 2169-3536. - 14:(2026 Jul 01), pp. 97841-97855. [10.1109/access.2026.3707213]

RIT-PDFMal-2026: A Comprehensive Benchmark Dataset for PDF Malware Detection

E. Damiani
Ultimo
2026

Abstract

Portable Document Format (PDF) has gained rapid popularity in the past two decades due to its platform-independence, and portability. A ingle PDF file can contain all the elements needed to convey a document in a fixed format, such as text, images, and fonts. This popularity made PDF documents an important attack vector for malicious activity. In this paper, we present a carefully curated comprehensive benchmark dataset for PDF malware detection. The dataset capture a large number of real-life malicious samples, and benign samples to build a 24,337-sample dataset suitable for training and testing of machine learning-based malicious PDF detection systems. The work also presents an extended set of features extracted statically from PDF files to support the detection process. The presented dataset was used in training and testing of a pipeline of different machine learning classifiers. Testing results showed that the extreme gradioent-boost classifeir was capable of detecting malicious PDF with F1 score exceeding 0.99. The model also performed similarly well in 10-fold cross validation. Additionally, the model was explained using Shapley additive explanation to ensure the model’s transparency and increase confidence and ensure the alignment of the classifier decision with human experience.
detection; machine learning; malware; pdf;
Settore INFO-01/A - Informatica
1-lug-2026
25-giu-2026
Article (author)
File in questo prodotto:
File Dimensione Formato  
RIT-PDFMal-2026_A_Comprehensive_Benchmark_Dataset_for_PDF_Malware_Detection(1).pdf

accesso aperto

Tipologia: Publisher's version/PDF
Licenza: Creative commons
Dimensione 2.61 MB
Formato Adobe PDF
2.61 MB Adobe PDF Visualizza/Apri
Pubblicazioni consigliate

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/2434/1258395
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 0
  • ???jsp.display-item.citation.isi??? ND
  • OpenAlex 0
social impact