JobbSafariLediga jobbMaster thesis - Vision Foundation Models enhanced Vision-Language Model 3D Spatiotemporal Reasoning

Master thesis - Vision Foundation Models enhanced Vision-Language Model 3D Spatiotemporal Reasoning
Scania GroupRekommenderat
Sammanfattning
This Master's thesis opportunity focuses on enhancing Vision-Language Models (VLMs) for autonomous driving by integrating representations from Vision Foundation Models (VFMs). The project, based in Södertälje, Sweden, involves exploring methods to improve spatiotemporal reasoning capabilities, with potential for scientific publication. Students will work with advanced deep learning architectures and access large-scale datasets and GPU resources, collaborating closely with researchers in the AI-4Jobbet i korthet
Arbetstid
deltid
Det här erbjuder vi
Opportunity to work with state-of-the-art deep learning architectures.Access to large-scale autonomous driving datasets and premium GPU computing infrastructure.Collaboration with researchers on next-generation AI for autonomous vehicles.Potential for contribution to scientific publications.
Ansök senast: Öppet tillsvidare
Publicerad: 2026-10-01
Beskrivning
30 credits - Using Vision Foundation Models to Improve Vision-Language Model 3D Spatiotemporal Reasoning for Autonomous Driving
Introduction
A Master's thesis is an excellent way to get closer to TRATON Group R&D and build relationships for the future.
Background
Vision-Language Models (VLMs) are increasingly being explored for autonomous driving and embodied AI due to their strong compositional and logical reasoning capabilities and generalizability. However, while these models excel at general visual understanding and reasoning, they often struggle with the fine-grained spatiotemporal perception required for safe and precise driving decisions. Conversely, modern Vision Foundation Models (VFMs) such as generative video world models (e.g., Wan2.1 [5], Cosmos [3]) and 3D geometry models (e.g., VGGT [6], DepthAnythingV2 [1]) have demonstrated exceptional ability in learning visual representations from large-scale data, capturing geometry and dynamics. This presents a highly promising opportunity by transferring the rich, low-level spatiotemporal priors from VFMs into the high-level reasoning frameworks of VLMs. However, the optimal strategies for combining these paradigms, without drastically increasing computational overhead and without deteriorating the language-aligned reasoning of the VLM (catastrophic forgetting), remain an open and exciting research challenge.
Objective
The objective of this thesis is to investigate methods for enhancing the spatiotemporal reasoning capabilities of VLMs by leveraging representations from state-of-the-art VFMs. The exact research questions will be formulated together with the student based on current literature, the student's interests, and the specific project direction. Possible research directions can include:
The Project Offers
Who are we looking for?
We are looking for a Master's student in Computer Science, Robotics, Engineering Physics, Electrical Engineering, Applied Mathematics, or a related field.
Experience with one or more of the following is beneficial:
The planned thesis start is January 2027.
Number of students: 1
Start date for the thesis work: [To be agreed]
Estimated time required: 20 weeks, full time (30 credits)
Contact persons and supervisors
Jesper Eriksson, jesper.ericsson@scania.com; Thomas Gustafsson, thomas.gustafsson@scania.com; Mohammad Nazari, mohammad.nazari@scania.com;
Hiring Manager: Maria Linnarsson, maria.linnarsson@scania.com
Application
Your application must include a CV, personal letter, and transcript of grades.
A background check might be conducted for this position. We are conducting interviews continuously and may close the recruitment earlier than the date specified.
References
Publication date:
1.10.2026 - 30.11.2026 (applications evaluated continuously)
Requisition ID: 33772
Number of Openings: 1.0
Part-time / Full-time: Full-time
Permanent / Temporary: Temporary
Country/Region: SE
Location(s):
Södertälje, SE, 151 38
Required Travel: 0%
Workplace: Hybrid
Introduction
A Master's thesis is an excellent way to get closer to TRATON Group R&D and build relationships for the future.
Background
Vision-Language Models (VLMs) are increasingly being explored for autonomous driving and embodied AI due to their strong compositional and logical reasoning capabilities and generalizability. However, while these models excel at general visual understanding and reasoning, they often struggle with the fine-grained spatiotemporal perception required for safe and precise driving decisions. Conversely, modern Vision Foundation Models (VFMs) such as generative video world models (e.g., Wan2.1 [5], Cosmos [3]) and 3D geometry models (e.g., VGGT [6], DepthAnythingV2 [1]) have demonstrated exceptional ability in learning visual representations from large-scale data, capturing geometry and dynamics. This presents a highly promising opportunity by transferring the rich, low-level spatiotemporal priors from VFMs into the high-level reasoning frameworks of VLMs. However, the optimal strategies for combining these paradigms, without drastically increasing computational overhead and without deteriorating the language-aligned reasoning of the VLM (catastrophic forgetting), remain an open and exciting research challenge.
Objective
The objective of this thesis is to investigate methods for enhancing the spatiotemporal reasoning capabilities of VLMs by leveraging representations from state-of-the-art VFMs. The exact research questions will be formulated together with the student based on current literature, the student's interests, and the specific project direction. Possible research directions can include:
- Exploring distillation methods to transfer spatiotemporal knowledge from VFMs to VLMs [2, 4].
- Investigating architectural modifications to e.g., insert or fuse VFM features into the VLM pipeline [7, 8].
- Evaluating the enhanced model on autonomous driving benchmarks that require complex temporal and spatial reasoning (e.g., spatial Vision-Question-Answering (VQA), 3D perception-based tasks, driving trajectory prediction).
The Project Offers
- The opportunity to work with state-of-the-art deep learning architectures at the intersection of vision and language.
- Access to large-scale autonomous driving datasets and premium GPU computing infrastructure (e.g., A100 40-80Gb clusters).
- Close collaboration with researchers working on next-generation AI for autonomous vehicles.
- A chance to contribute to an active research area with the potential for scientific publication.
Who are we looking for?
We are looking for a Master's student in Computer Science, Robotics, Engineering Physics, Electrical Engineering, Applied Mathematics, or a related field.
Experience with one or more of the following is beneficial:
- Deep learning and machine learning
- Computer vision and/or natural language processing
- PyTorch or similar frameworks
- Autonomous driving or robotics
The planned thesis start is January 2027.
Number of students: 1
Start date for the thesis work: [To be agreed]
Estimated time required: 20 weeks, full time (30 credits)
Contact persons and supervisors
Jesper Eriksson, jesper.ericsson@scania.com; Thomas Gustafsson, thomas.gustafsson@scania.com; Mohammad Nazari, mohammad.nazari@scania.com;
Hiring Manager: Maria Linnarsson, maria.linnarsson@scania.com
Application
Your application must include a CV, personal letter, and transcript of grades.
A background check might be conducted for this position. We are conducting interviews continuously and may close the recruitment earlier than the date specified.
References
- Haotong Lin, Sili Chen, Jun Hao Liew, Donny Y. Chen, Zhenyu Li, Guang Shi, Jiashi Feng, and Bingyi Kang. Depth Anything 3: Recovering the visual space from any views. arXiv preprint arXiv:2511.10647, 2025
- Yuechen Luo, Fang Li, Shaoqing Xu, Yang Ji, Zehan Zhang, Bing Wang, Yuannan Shen, Jianwei Cui, Long Chen, Guang Chen, Hangjun Ye, Zhi-Xin Yang, and Fuxi Wen. LaST-VLA: Thinking in latent spatio-temporal space for vision-language-action in autonomous driving. arXiv preprint arXiv:2603.01928, 2026
- NVIDIA, Niket Agarwal, Arslan Ali, Maciej Bala, Yogesh Balaji, Erik Barker, et al. Cosmos world foundation model platform for physical AI. arXiv preprint arXiv:2501.03575, 2025
- Yiming Qin, Bomin Wei, Jiaxin Ge, Konstantinos Kallidromitis, Stephanie Fu, Trevor Darrell, and XuDong Wang. Chain-of-Visual-Thought: Teaching VLMs to see and think better with continuous visual tokens. arXiv preprint arXiv:2511.19418, 2025
- Team Wan, Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, et al. Wan: Open and advanced large-scale video generative models. arXiv preprint arXiv:2503.20314, 2025
- Jianyuan Wang, Minghao Chen, Shangzhan Zhang, Nikita Karaev, Johannes Schönberger, Patrick Labatut, Piotr Bojanowski, David Novotny, Andrea Vedaldi, and Christian Rupprecht. VGGT-Ω. arXiv preprint arXiv:2605.15195, 2026
- Jie Wang, Guang Li, Zhijian Huang, Chenxu Dang, Hangjun Ye, Yahong Han, and Long Chen. VGGDrive: Empowering vision-language models with cross-view geometric grounding for autonomous driving. arXiv preprint arXiv:2602.20794, 2026
- Jin Yao, Dhruva Dixith Kurra, Tom Lampo, Zezhou Cheng, Danhua Guo, and Burhan Yaman. VLGA: Vision-language-geometry-action models for autonomous driving. arXiv preprint arXiv:2606.12396, 2026
Publication date:
1.10.2026 - 30.11.2026 (applications evaluated continuously)
Requisition ID: 33772
Number of Openings: 1.0
Part-time / Full-time: Full-time
Permanent / Temporary: Temporary
Country/Region: SE
Location(s):
Södertälje, SE, 151 38
Required Travel: 0%
Workplace: Hybrid
Ansök till tjänsten
Master thesis - Vision Foundation Models enhanced Vision-Language Model 3D Spatiotemporal Reasoning
Rekommenderat
OM FÖRETAGET

Scania Group








