Portable Document Format (PDF) has gained rapid popularity in the past two decades due to its platform-independence, and portability. A ingle PDF file can contain all the elements needed to convey a document in a fixed format, such as text, images, and fonts. This popularity made PDF documents an important attack vector for malicious activity. In this paper, we present a carefully curated comprehensive benchmark dataset for PDF malware detection. The dataset capture a large number of real-life malicious samples, and benign samples to build a 24,337-sample dataset suitable for training and testing of machine learning-based malicious PDF detection systems. The work also presents an extended set of features extracted statically from PDF files to support the detection process. The presented dataset was used in training and testing of a pipeline of different machine learning classifiers. Testing results showed that the extreme gradioent-boost classifeir was capable of detecting malicious PDF with F1 score exceeding 0.99. The model also performed similarly well in 10-fold cross validation. Additionally, the model was explained using Shapley additive explanation to ensure the model’s transparency and increase confidence and ensure the alignment of the classifier decision with human experience.
RIT-PDFMal-2026: A Comprehensive Benchmark Dataset for PDF Malware Detection / M.M. Alani, E.D.. - In: IEEE ACCESS. - ISSN 2169-3536. - 14:(2026 Jul 01), pp. 97841-97855. [10.1109/access.2026.3707213]
RIT-PDFMal-2026: A Comprehensive Benchmark Dataset for PDF Malware Detection
E. DamianiUltimo
2026
Abstract
Portable Document Format (PDF) has gained rapid popularity in the past two decades due to its platform-independence, and portability. A ingle PDF file can contain all the elements needed to convey a document in a fixed format, such as text, images, and fonts. This popularity made PDF documents an important attack vector for malicious activity. In this paper, we present a carefully curated comprehensive benchmark dataset for PDF malware detection. The dataset capture a large number of real-life malicious samples, and benign samples to build a 24,337-sample dataset suitable for training and testing of machine learning-based malicious PDF detection systems. The work also presents an extended set of features extracted statically from PDF files to support the detection process. The presented dataset was used in training and testing of a pipeline of different machine learning classifiers. Testing results showed that the extreme gradioent-boost classifeir was capable of detecting malicious PDF with F1 score exceeding 0.99. The model also performed similarly well in 10-fold cross validation. Additionally, the model was explained using Shapley additive explanation to ensure the model’s transparency and increase confidence and ensure the alignment of the classifier decision with human experience.| File | Dimensione | Formato | |
|---|---|---|---|
|
RIT-PDFMal-2026_A_Comprehensive_Benchmark_Dataset_for_PDF_Malware_Detection(1).pdf
accesso aperto
Tipologia:
Publisher's version/PDF
Licenza:
Creative commons
Dimensione
2.61 MB
Formato
Adobe PDF
|
2.61 MB | Adobe PDF | Visualizza/Apri |
Pubblicazioni consigliate
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.




