JobbSafariLediga jobbMaster thesis - LiDAR-Enhanced Vision-Language-Action Models for Safer Autonomous Truck

Master thesis - LiDAR-Enhanced Vision-Language-Action Models for Safer Autonomous Truck
Scania GroupRekommenderat
Sammanfattning
The MOPS project is seeking a Master's student to explore the integration of LiDAR data into NVIDIA's Alpamayo models to enhance the safety of autonomous trucks. This research will focus on improving spatial reasoning and trajectory prediction in complex driving scenarios. The position is based in Södertälje, Sweden, and involves a full-time commitment for 20 weeks, starting in January 2027. The student will work closely with industry and academic partners, contributing to significant research.Jobbet i korthet
Arbetstid
deltid
Det här erbjuder vi
Supervision from TRATON and collaboration with academic and industrial partners.Opportunity to contribute to a significant industrial research project.Development of a reusable LiDAR-enhanced model for future research.
Ansök senast: Öppet tillsvidare
Publicerad: 2026-10-01
Beskrivning
30 credits - LiDAR-Enhanced Vision-Language-Action Models for Safer Autonomous Trucks
Introduction
The VINNOVA-FFI Multi-modal Open-set Perception for Safer Autonomous Trucks (MOPS) project investigates how multimodal sensing and open-set perception can improve the safety of autonomous heavy-duty vehicles. Autonomous trucks must operate in open-world environments containing rare, unfamiliar, or previously unseen objects and situations. Their size, braking distance, limited maneuverability, and operational environments place particularly high demands on reliable perception and spatial reasoning.
Background
Camera-based end-to-end driving models capture rich semantic information, but their estimates of geometry, depth, and object distance may be sensitive to poor visibility, occlusion, and visually ambiguous scenes. LiDAR provides complementary three-dimensional measurements that could improve geometric understanding and help a driving model respond safely to unfamiliar objects and scenarios.
NVIDIA Alpamayo is a family of open vision-language-action models for autonomous driving. Its models process multi-camera observations and driving context to generate trajectories together with interpretable Chain-of-Causation reasoning traces. Current open releases primarily rely on camera video and ego-motion inputs, creating an opportunity to investigate how LiDAR can be incorporated into the model's perception, reasoning, and action-prediction pipeline.
The thesis will initially rely on open-source Alpamayo models and training recipes and suitable datasets containing synchronized camera, LiDAR, and vehicle-motion data, such as the NVIDIA Physical AI Autonomous Vehicles dataset.
Objective
The thesis will investigate how LiDAR information can be integrated into Alpamayo to improve spatial reasoning, trajectory prediction, and robustness to unfamiliar objects and situations.
Possible research directions include:
The proposed multimodal model should be compared with a camera-only Alpamayo baseline. Evaluation may include trajectory accuracy, collision-related metrics, open-set detection or response metrics, reasoning-action consistency, robustness, and computational efficiency. Where suitable data and benchmarks are available, the study may examine scenarios particularly relevant to heavy-duty vehicles, such as long stopping distances, narrow clearances, long-range perception, unusual road users, and unfamiliar obstacles.
The exact model variant, fusion strategy, dataset, and evaluation framework will be finalized together with the selected student and the MOPS project team, based on available open-source tools, computational resources, and the project's priorities.
The Project Offers
The student will contribute to the MOPS industrial research project and work on a problem combining multimodal machine learning, autonomous driving, three-dimensional perception, open-set recognition, and vision-language-action modeling. The work is expected to produce a reusable LiDAR-enhanced Alpamayo baseline, sensor-fusion component, or experimental framework that can support continued MOPS research.
The student will receive supervision from TRATON and have opportunities to interact with the project's academic and industrial collaborators.
Who are we looking for?
We are looking for a Master's student in computer science, machine learning, robotics, engineering physics, electrical engineering, or a related field.
A suitable candidate should have:
The planned thesis start is January 2027.
Number of students: 1
Start date for the thesis work: [To be agreed]
Estimated time required: 20 weeks, full time (30 credits)
Contact persons and supervisors
Thomas Gustafsson, thomas.gustafsson@scania.com; Jesper Eriksson, jesper.ericsson@scania.com;
Hiring Manager: Maria Linnarsson, maria.linnarsson@scania.com
Application
Your application must include a CV, personal letter, and transcript of grades.
A background check might be conducted for this position. We are conducting interviews continuously and may close the recruitment earlier than the date specified.
Publication date:
1.10.2026 - 30.11.2026 (applications evaluated continuously)
Requisition ID: 33769
Number of Openings: 1.0
Part-time / Full-time: Full-time
Permanent / Temporary: Temporary
Country/Region: SE
Location(s):
Södertälje, SE, 151 38
Required Travel: 0%
Workplace: Hybrid
Introduction
The VINNOVA-FFI Multi-modal Open-set Perception for Safer Autonomous Trucks (MOPS) project investigates how multimodal sensing and open-set perception can improve the safety of autonomous heavy-duty vehicles. Autonomous trucks must operate in open-world environments containing rare, unfamiliar, or previously unseen objects and situations. Their size, braking distance, limited maneuverability, and operational environments place particularly high demands on reliable perception and spatial reasoning.
Background
Camera-based end-to-end driving models capture rich semantic information, but their estimates of geometry, depth, and object distance may be sensitive to poor visibility, occlusion, and visually ambiguous scenes. LiDAR provides complementary three-dimensional measurements that could improve geometric understanding and help a driving model respond safely to unfamiliar objects and scenarios.
NVIDIA Alpamayo is a family of open vision-language-action models for autonomous driving. Its models process multi-camera observations and driving context to generate trajectories together with interpretable Chain-of-Causation reasoning traces. Current open releases primarily rely on camera video and ego-motion inputs, creating an opportunity to investigate how LiDAR can be incorporated into the model's perception, reasoning, and action-prediction pipeline.
The thesis will initially rely on open-source Alpamayo models and training recipes and suitable datasets containing synchronized camera, LiDAR, and vehicle-motion data, such as the NVIDIA Physical AI Autonomous Vehicles dataset.
Objective
The thesis will investigate how LiDAR information can be integrated into Alpamayo to improve spatial reasoning, trajectory prediction, and robustness to unfamiliar objects and situations.
Possible research directions include:
- representing LiDAR data as point-cloud, voxel, range-image, bird's-eye-view, or learned tokens suitable for a vision-language-action model;
- comparing early, intermediate, and late fusion of camera and LiDAR features;
- developing token-efficient fusion mechanisms that limit the additional memory and inference cost;
- fine-tuning a suitable open Alpamayo release using synchronized camera and LiDAR observations;
- evaluating whether LiDAR improves trajectory prediction and spatial reasoning around previously unseen objects;
- studying whether the model's reasoning traces correctly identify uncertainty, unfamiliar objects, and safety-relevant geometric constraints;
- evaluating robustness under poor visibility, occlusion, sensor degradation, or missing modalities; or
- analyzing which truck-relevant driving scenarios benefit most from explicit three-dimensional information.
The proposed multimodal model should be compared with a camera-only Alpamayo baseline. Evaluation may include trajectory accuracy, collision-related metrics, open-set detection or response metrics, reasoning-action consistency, robustness, and computational efficiency. Where suitable data and benchmarks are available, the study may examine scenarios particularly relevant to heavy-duty vehicles, such as long stopping distances, narrow clearances, long-range perception, unusual road users, and unfamiliar obstacles.
The exact model variant, fusion strategy, dataset, and evaluation framework will be finalized together with the selected student and the MOPS project team, based on available open-source tools, computational resources, and the project's priorities.
The Project Offers
The student will contribute to the MOPS industrial research project and work on a problem combining multimodal machine learning, autonomous driving, three-dimensional perception, open-set recognition, and vision-language-action modeling. The work is expected to produce a reusable LiDAR-enhanced Alpamayo baseline, sensor-fusion component, or experimental framework that can support continued MOPS research.
The student will receive supervision from TRATON and have opportunities to interact with the project's academic and industrial collaborators.
Who are we looking for?
We are looking for a Master's student in computer science, machine learning, robotics, engineering physics, electrical engineering, or a related field.
A suitable candidate should have:
- strong programming skills, preferably in Python and PyTorch;
- knowledge of machine learning and deep learning;
- an interest in autonomous driving, multimodal transformers, computer vision, or three-dimensional perception;
- experience with LiDAR, point-cloud processing, open-set recognition, or large-scale model fine-tuning is beneficial but not required; and
- motivation to combine scientific investigation with practical implementation.
The planned thesis start is January 2027.
Number of students: 1
Start date for the thesis work: [To be agreed]
Estimated time required: 20 weeks, full time (30 credits)
Contact persons and supervisors
Thomas Gustafsson, thomas.gustafsson@scania.com; Jesper Eriksson, jesper.ericsson@scania.com;
Hiring Manager: Maria Linnarsson, maria.linnarsson@scania.com
Application
Your application must include a CV, personal letter, and transcript of grades.
A background check might be conducted for this position. We are conducting interviews continuously and may close the recruitment earlier than the date specified.
Publication date:
1.10.2026 - 30.11.2026 (applications evaluated continuously)
Requisition ID: 33769
Number of Openings: 1.0
Part-time / Full-time: Full-time
Permanent / Temporary: Temporary
Country/Region: SE
Location(s):
Södertälje, SE, 151 38
Required Travel: 0%
Workplace: Hybrid
Ansök till tjänsten
Master thesis - LiDAR-Enhanced Vision-Language-Action Models for Safer Autonomous Truck
Rekommenderat
OM FÖRETAGET

Scania Group








