Accurate molecular toxicity prediction is essential in drug discovery and environmental safety. Traditional laboratory-based toxicity testing is costly and time-consuming, limiting its applicability for large-scale chemical screening. As a result, deep learning based approaches for toxicity prediction have emerged as an effective alternative. However, existing approaches struggle to capture the heterogeneous and multiscale nature of molecular structure. To address this, we propose MMAF (Multimodal Attention Fusion Network), a multimodal deep learning framework that integrates sequential, structural, and substructure-level information. More specifically, we encode SMILES sequences, molecular graphs, and fingerprint-derived substructure information using modality-specific encoders. Then, a novel interaction-aware fusion mechanism models cross-modal dependencies and integrates diverse molecular views into a unified representation. Experiments across six MoleculeNet benchmark datasets show that MMAF consistently outperforms single and dual-modality baselines. The code is available at https://github.com/CVML-CFU/MMAF.
MMAF: Multimodal Attention Fusion for Molecular Toxicity Prediction
Rehman, Faiz Ur;Rahman, Muhammad Rameez Ur;Vascon, Sebastiano;Pelillo, Marcello
2026
Abstract
Accurate molecular toxicity prediction is essential in drug discovery and environmental safety. Traditional laboratory-based toxicity testing is costly and time-consuming, limiting its applicability for large-scale chemical screening. As a result, deep learning based approaches for toxicity prediction have emerged as an effective alternative. However, existing approaches struggle to capture the heterogeneous and multiscale nature of molecular structure. To address this, we propose MMAF (Multimodal Attention Fusion Network), a multimodal deep learning framework that integrates sequential, structural, and substructure-level information. More specifically, we encode SMILES sequences, molecular graphs, and fingerprint-derived substructure information using modality-specific encoders. Then, a novel interaction-aware fusion mechanism models cross-modal dependencies and integrates diverse molecular views into a unified representation. Experiments across six MoleculeNet benchmark datasets show that MMAF consistently outperforms single and dual-modality baselines. The code is available at https://github.com/CVML-CFU/MMAF.I documenti in ARCA sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.



