Scania
780 open roles
Master thesis - Vision Foundation Models enhanced Vision-Language Model 3D Spatiotemporal Reasoning
Södertälje, SE, 151 38Posted Oct 10, 2026
Job description
30 credits
- Using Vision Foundation Models to Improve Vision-Language Model 3D Spatiotemporal Reasoning for Autonomous Driving Introduction A Master's thesis is an excellent way to get closer to TRATON Group R&D and build relationships for the future. Background Vision-Language Models (VLMs) are increasingly being explored for autonomous driving and embodied AI due to their strong compositional and logical reasoning capabilities and generalizability. However, while these models excel at general visual understanding and reasoning, they often struggle with the fine-grained spatiotemporal perception required for safe and precise driving decisions. Conversely, modern Vision Foundation Models (VFMs) such as generative video world models (e.g., Wan2.1 [5], Cosmos [3]) and 3D geometry models (e.g., VGGT [6], Depth. AnythingV2 [1]) have demonstrated exceptional ability in learning visual representations from large-scale data, capturing geometry and dynamics. This presents a highly promising opportunity by transferring the rich, low-level spatiotemporal priors from VFMs into the high-level reasoning frameworks of VLMs. However, the optimal strategies for combining these paradigms, without drastically increasing computational overhead and without deteriorating the language-aligned reasoning of the VLM (catastrophic forgetting), remain an open and exciting research challenge. Objective The objective of this thesis is to investigate methods for enhancing the spatiotemporal reasoning capabilities of VLMs by leveraging representations from state-of-the-art VFMs. The exact research questions will be formulated together with the student based on current literature, the student’s interests, and the specific project direction. Possible research directions can include: Exploring distillation methods to transfer spatiotemporal knowledge from VFMs to VLMs [2, 4]. Investigating architectural modifications to e.g., insert or fuse VFM features into the VLM pipeline [7, 8]. Evaluating the enhanced model on autonomous driving benchmarks that require complex temporal and spatial reasoning (e.g., spatial Vision-Question-Answering (VQA), 3D perception-based tasks, driving trajectory prediction). The Project Offers The opportunity to work with state-of-the-art deep learning architectures at the intersection of vision and language. Access to large-scale autonomous driving datasets and premium GPU computing infrastructure (e.g., A100 40-80Gb clusters). Close collaboration with researchers working on next-generation AI for autonomous vehicles. A chance to contribute to an active research area with the potential for scientific publication. Who are we looking for? We are looking for a Master’s student in Computer Science, Robotics, Engineering Physics, Electrical Engineering, Applied Mathematics, or a related field. Experience with one or more of the following is beneficial: Deep learning and machine learning Computer vision and/or natural language processing Py. Torch or similar frameworks Autonomous driving or robotics The planned thesis start is January 2027. Number of students: 1 Start date for the thesis work: [To be agreed] Estimated time required: 20 weeks, full time (30 credits) Contact persons and supervisors Jesper Eriksson, jesper.ericsson@scania.com ; Thomas Gustafsson, thomas.gustafsson@scania.com ; Mohammad Nazari, mohammad.nazari@scania.com ; Hiring Manager: Maria Linnarsson, maria.linnarsson@scania.com Application Your application must include a CV, personal letter, and transcript of grades. A background check might be conducted for this position. We are conducting interviews continuously and may close the recruitment earlier than the date specified. References Haotong Lin, Sili Chen, Jun Hao Liew, Donny Y. Chen, Zhenyu Li, Guang Shi, Jiashi Feng, and Bingyi Kang. Depth Anything 3: Recovering the visual space from any views. ar. Xiv preprint ar. Xiv:2511.10647, 2025 Yuechen Luo, Fang Li, Shaoqing Xu, Yang Ji, Zehan Zhang, Bing Wang, Yuannan Shen, Jianwei Cui, Long Chen, Guang Chen, Hangjun Ye, Zhi-Xin Yang, and Fuxi Wen. LaST-VLA: Thinking in latent spatio-temporal space for vision-language-action in autonomous driving. ar. Xiv preprint ar. Xiv:2603.01928, 2026 NVIDIA, Niket Agarwal, Arslan Ali, Maciej Bala, Yogesh Balaji, Erik Barker, et al. Cosmos world foundation model platform for physical AI. ar. Xiv preprint ar. Xiv:2501.03575, 2025 Yiming Qin, Bomin Wei, Jiaxin Ge, Konstantinos Kallidromitis, Stephanie Fu, Trevor Darrell, and Xu. Dong Wang. Chain-of-Visual-Thought: Teaching VLMs to see and think better with continuous visual tokens. ar. Xiv preprint ar. Xiv:2511.19418, 2025 Team Wan, Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, et al. Wan: Open and advanced large-scale video generative models. ar. Xiv preprint ar. Xiv:2503.20314, 2025 Jianyuan Wang, Minghao Chen, Shangzhan Zhang, Nikita Karaev, Johannes Schönberger, Patrick Labatut, Piotr Bojanowski, David Novotny, Andrea Vedaldi, and Christian Rupprecht. VGGT-Ω. ar. Xiv preprint ar. Xiv:2605.15195, 2026 Jie Wang, Guang Li, Zhijian Huang, Chenxu Dang, Hangjun Ye, Yahong Han, and Long Chen. VGGDrive: Empowering vision-language models with cross-view geometric grounding for autonomous driving. ar. Xiv preprint ar. Xiv:2602.20794, 2026 Jin Yao, Dhruva Dixith Kurra, Tom Lampo, Zezhou Cheng, Danhua Guo, and Burhan Yaman. VLGA: Vision-language-geometry-action models for autonomous driving. ar. Xiv preprint ar. Xiv:2606.12396, 2026 Publication date: 1.10.2026 - 30.11.2026 (applications evaluated continuously)
Description copied from Scania's careers page. Read the full posting before you apply.
More jobs at Scania
UH-tekniker med inriktning El och automation
Scania· Södertälje, SE, 151 38Truckförare/Montörer till Scania Oskarshamn
Scania· Oskarshamn, SE, 572 36Automehāniķis
Scania· Riga, LV, LV-1058Mechatroniker (m/w/d) für München/Oberschleißheim
Scania· Hicklstr. 4, 85764 Oberschleiß, DEAutomehāniķis
Scania· Riga, LV, LV-1058