Sök med AI
JobbSafariLediga jobbMaster thesis - Vision Foundation Models enhanced Vision-Language Model 3D Spatiotemporal Reasoning

Master thesis - Vision Foundation Models enhanced Vision-Language Model 3D Spatiotemporal Reasoning

Scania Group
Rekommenderat

Sammanfattning

This Master's thesis opportunity focuses on enhancing Vision-Language Models (VLMs) for autonomous driving by integrating representations from Vision Foundation Models (VFMs). The project, based in Södertälje, Sweden, involves exploring methods to improve spatiotemporal reasoning capabilities, with potential for scientific publication. Students will work with advanced deep learning architectures and access large-scale datasets and GPU resources, collaborating closely with researchers in the AI-4
Visa hela jobbannonsen

Jobbet i korthet

Arbetstid

deltid


Det här erbjuder vi

Opportunity to work with state-of-the-art deep learning architectures.Access to large-scale autonomous driving datasets and premium GPU computing infrastructure.Collaboration with researchers on next-generation AI for autonomous vehicles.Potential for contribution to scientific publications.

Södertälje

Ansök senast: Öppet tillsvidare
Publicerad: 2026-10-01

Beskrivning

30 credits - Using Vision Foundation Models to Improve Vision-Language Model 3D Spatiotemporal Reasoning for Autonomous Driving


Introduction


A Master's thesis is an excellent way to get closer to TRATON Group R&D and build relationships for the future.

Background



Vision-Language Models (VLMs) are increasingly being explored for autonomous driving and embodied AI due to their strong compositional and logical reasoning capabilities and generalizability. However, while these models excel at general visual understanding and reasoning, they often struggle with the fine-grained spatiotemporal perception required for safe and precise driving decisions. Conversely, modern Vision Foundation Models (VFMs) such as generative video world models (e.g., Wan2.1 [5], Cosmos [3]) and 3D geometry models (e.g., VGGT [6], DepthAnythingV2 [1]) have demonstrated exceptional ability in learning visual representations from large-scale data, capturing geometry and dynamics. This presents a highly promising opportunity by transferring the rich, low-level spatiotemporal priors from VFMs into the high-level reasoning frameworks of VLMs. However, the optimal strategies for combining these paradigms, without drastically increasing computational overhead and without deteriorating the language-aligned reasoning of the VLM (catastrophic forgetting), remain an open and exciting research challenge.

Objective



The objective of this thesis is to investigate methods for enhancing the spatiotemporal reasoning capabilities of VLMs by leveraging representations from state-of-the-art VFMs. The exact research questions will be formulated together with the student based on current literature, the student's interests, and the specific project direction. Possible research directions can include:

  • Exploring distillation methods to transfer spatiotemporal knowledge from VFMs to VLMs [2, 4].


  • Investigating architectural modifications to e.g., insert or fuse VFM features into the VLM pipeline [7, 8].


  • Evaluating the enhanced model on autonomous driving benchmarks that require complex temporal and spatial reasoning (e.g., spatial Vision-Question-Answering (VQA), 3D perception-based tasks, driving trajectory prediction).


The Project Offers



  • The opportunity to work with state-of-the-art deep learning architectures at the intersection of vision and language.


  • Access to large-scale autonomous driving datasets and premium GPU computing infrastructure (e.g., A100 40-80Gb clusters).


  • Close collaboration with researchers working on next-generation AI for autonomous vehicles.


  • A chance to contribute to an active research area with the potential for scientific publication.


Who are we looking for?



We are looking for a Master's student in Computer Science, Robotics, Engineering Physics, Electrical Engineering, Applied Mathematics, or a related field.

Experience with one or more of the following is beneficial:

  • Deep learning and machine learning


  • Computer vision and/or natural language processing


  • PyTorch or similar frameworks


  • Autonomous driving or robotics


The planned thesis start is January 2027.

Number of students: 1

Start date for the thesis work: [To be agreed]

Estimated time required: 20 weeks, full time (30 credits)

Contact persons and supervisors



Jesper Eriksson, jesper.ericsson@scania.com; Thomas Gustafsson, thomas.gustafsson@scania.com; Mohammad Nazari, mohammad.nazari@scania.com;

Hiring Manager: Maria Linnarsson, maria.linnarsson@scania.com

Application



Your application must include a CV, personal letter, and transcript of grades.

A background check might be conducted for this position. We are conducting interviews continuously and may close the recruitment earlier than the date specified.

References

  1. Haotong Lin, Sili Chen, Jun Hao Liew, Donny Y. Chen, Zhenyu Li, Guang Shi, Jiashi Feng, and Bingyi Kang. Depth Anything 3: Recovering the visual space from any views. arXiv preprint arXiv:2511.10647, 2025


  1. Yuechen Luo, Fang Li, Shaoqing Xu, Yang Ji, Zehan Zhang, Bing Wang, Yuannan Shen, Jianwei Cui, Long Chen, Guang Chen, Hangjun Ye, Zhi-Xin Yang, and Fuxi Wen. LaST-VLA: Thinking in latent spatio-temporal space for vision-language-action in autonomous driving. arXiv preprint arXiv:2603.01928, 2026


  1. NVIDIA, Niket Agarwal, Arslan Ali, Maciej Bala, Yogesh Balaji, Erik Barker, et al. Cosmos world foundation model platform for physical AI. arXiv preprint arXiv:2501.03575, 2025


  1. Yiming Qin, Bomin Wei, Jiaxin Ge, Konstantinos Kallidromitis, Stephanie Fu, Trevor Darrell, and XuDong Wang. Chain-of-Visual-Thought: Teaching VLMs to see and think better with continuous visual tokens. arXiv preprint arXiv:2511.19418, 2025


  1. Team Wan, Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, et al. Wan: Open and advanced large-scale video generative models. arXiv preprint arXiv:2503.20314, 2025


  1. Jianyuan Wang, Minghao Chen, Shangzhan Zhang, Nikita Karaev, Johannes Schönberger, Patrick Labatut, Piotr Bojanowski, David Novotny, Andrea Vedaldi, and Christian Rupprecht. VGGT-Ω. arXiv preprint arXiv:2605.15195, 2026


  1. Jie Wang, Guang Li, Zhijian Huang, Chenxu Dang, Hangjun Ye, Yahong Han, and Long Chen. VGGDrive: Empowering vision-language models with cross-view geometric grounding for autonomous driving. arXiv preprint arXiv:2602.20794, 2026


  1. Jin Yao, Dhruva Dixith Kurra, Tom Lampo, Zezhou Cheng, Danhua Guo, and Burhan Yaman. VLGA: Vision-language-geometry-action models for autonomous driving. arXiv preprint arXiv:2606.12396, 2026


Publication date:

1.10.2026 - 30.11.2026 (applications evaluated continuously)

Requisition ID: 33772

Number of Openings: 1.0

Part-time / Full-time: Full-time

Permanent / Temporary: Temporary

Country/Region: SE

Location(s):
Södertälje, SE, 151 38

Required Travel: 0%

Workplace: Hybrid

Ansök till tjänsten

Master thesis - Vision Foundation Models enhanced Vision-Language Model 3D Spatiotemporal Reasoning

Rekommenderat
Denna arbetsplats har annonserats på Scania-tjänsten den 2026-10-01 och publicerades av Scania.
Tillbaka till toppen

OM FÖRETAGET

Scania Group
Visa alla jobb för Scania Group

Hittade du inte vad du letade efter?

Beskriv med dina egna ord vad du söker, precis som om du skulle förklara det för en kompis. Josi hittar jobb som matchar dig på riktigt.
Testa nu

Sök efter fler liknande jobb

SödertäljeForskning och utvecklingThesis workBachelor's ThesisMaster's Thesis

Läs också

Uppdämda jobbdrömmar: Svenskarna vill vidare men marknaden står still
För arbetsgivare

Uppdämda jobbdrömmar: Svenskarna vill vidare men marknaden står still

Antalet jobbannonser i Sverige ökade i juni för andra månaden i rad, vilket visar en fortsatt positiv trend på arbetsmarknaden.

Lästid 3 min

Liknande jobb

Visa alla lediga jobb
Scania Group

Master thesis - Evaluating and Improving Explainable Vision-Language Models for Autonomous driving

Södertälje
1/10 – tillsvidare

Jobb per stad

Det är enklare än någonsin att söka jobb – men svårare än någonsin att hitta rätt. Det vill vi ändra på. JobbSafari är din guide genom arbetslivet, byggd för att matcha rätt person med rätt möjlighet bland tusentals lediga jobb i Sverige.

JobbSafari är en del av Duunitori Group – Duunitori är Finlands största jobbsökmotor och en betrodd partner inom rekrytering, rekryteringsmarknadsföring och employer branding.

Stockholm, Sweden

JobbSafari AB

Grev Turegatan 11A

114 46 Stockholm, Sweden

info@jobbsafari.se

+46 (0) 8 515 10 774

Helsinki, Finland

Duunitori Oy

Toinen Linja 7

00530 Helsinki, Finland

asiakaspalvelu@duunitori.fi

+358 44 980 3558

Norway

Jobbland AS

c/o EMU Growth Partners Norway AS

Mercurveien 86

9408 Harstad

Norway

info@jobbsafari.se

+46 70 314 59 79

  • jobbsafari.se
  • duunitori.fi
  • jobbsafari.no
  • allaloner.se
  • jobbland.se