
Job description
The Video Computer Vision (VCV) organization is a centralized applied research and engineering team developing real-time, on-device Computer Vision and Machine Perception technologies across Apple products. Within VCV, the Multimodal Intelligence team builds next-generation multimodal AI systems that combine large language models, multimodal LLMs, and foundation models to create intelligent systems capable of understanding, reasoning, and acting across language, vision, audio, and tools.
We develop multimodal agentic systems deeply integrated into the Apple ecosystem, partnering across hardware, software, and ML teams to deliver advanced AI in real-time, scalable, and privacy-preserving experiences reaching millions of users.