Vision Transformer-Based Autism Detection
A Vision Transformer framework with multimodal fusion of facial and behavioral signals for early, interpretable Autism Spectrum Disorder screening in toddlers.
The idea
Autism Spectrum Disorder screening today mostly depends on clinician-administered behavioral assessments, which are slow to access and unevenly available. This project builds a Vision Transformer (ViT) framework that fuses facial imaging with behavioral signal data into a single multimodal model, aiming at scalable, low-cost early screening for ASD in toddlers that could run outside a clinical setting.
Status
The goal is an interpretable, low-friction screening tool rather than a diagnostic replacement — explainability matters here because a false negative or an opaque prediction in a pediatric health context has real consequences. The work is ongoing and has been submitted for journal publication; method and result details are held back until the review process concludes.
Tech stack & key skills
Core tools, methods and skills demonstrated in this project: