Senior Computer Vision Engineer, Embodied AI
Senior Computer Vision Engineer, Embodied AI at Niantic Spatial — San Francisco, CA, US
- Company: Niantic Spatial
- Location: San Francisco, CA, US
- Employment type: Full-time
- Salary: $229,500–$255,000 a year
- Posted: 2026-09-23
About this role
About Niantic Spatial
At Niantic Spatial, we're building the future of physical AI. Powered by a proprietary database of over 30 billion posed images, our groundbreaking mapping technology unlocks a new dimension of interaction and spatial intelligence that helps both humans and machines better understand, represent, navigate, and engage with the real environment.
Our reconstruction technology captures environments with geometric accuracy and extreme detail from any standard camera, and our Visual Positioning System delivers precise positioning almost anywhere in the world. We serve customers across robotics, the public sector, and energy and industrial markets - building for the 80% of economic activity that takes place beyond our screens.
About the Role
Niantic Spatial makes the physical world computable, helping people and machines collaborate safely by aligning how they understand reality. One of the fundamental problems we're helping autonomy teams and engineers overcome is the sim-to-real gap for visual-spatial understanding. We're applying our team's decades of experience encoding the world precisely as it is at Google Maps, Google Earth, and Niantic Labs to build the scalable real-to-sim stack for embodied AI.
We're looking for a Senior Computer Vision Engineer to join our Embodied AI team, focused on 3D scene understanding and semantics. A reconstructed environment is only useful to a robot once it knows what is in it: which surfaces are floor, which objects can be grasped, moved, or opened, where one room ends and another begins. You'll work hand in hand with our Research Scientists to take advances in open-vocabulary perception, 3D semantic labeling, and scene graph construction and turn them into reliable capabilities that embodied AI teams can use in production.
This is an applied research role. You'll live at the boundary between a research prototype and a production pipeline: adapting methods so they hold up on messy customer captures, lifting 2D semantic signals into geometrically consistent 3D, defining what "semantically correct" means for a downstream policy, and feeding what you learn back into the research agenda. You understand that the difficult part is often not the core method, but the messy inputs, edge cases, operational constraints, and downstream requirements surrounding it. You care about quality, cost, latency, and reliability, and you know that a system is only successful when people can depend on it. This is a hands-on engineering role for someone who wants to build, ship, and own outcomes. You'll help shape both the technology and the product, while remaining close to the code and the problems our customers are trying to solve.
What You'll Do
- Productionize 3D Scene Understanding Capabilities - Turn research prototypes in open-vocabulary segmentation, 3D semantic and instance labeling, affordance prediction, and scene graph construction into reliable pipeline components and services that operate across varied customer data and environments.
- Own End-to-End Output Quality - Ensure that reconstructed environments carry labels, instances, and relationships that are consistent, complete, and correct enough for downstream simulation and policy learning, and define what "correct enough" means for each customer workflow.
- Bridge Research and Production - Work day to day with the Embodied AI Research Scientist: pressure-test new methods on real customer data, identify where they break, and bring production and customer constraints back into research planning to prioritize the advances that matter most.
- Make Systems Robust to the Real World - Diagnose failures caused by capture quality, scene complexity, scale, calibration, coordinate conventions, and other assumptions that prototypes often leave implicit.
- Build Evaluation and Quality Infrastructure - Create datasets, regression tests, quality gates, and benchmarking tools that help the team measure whether changes improve the system.
- Improve Performance, Throughput, and Cost - Profile and optimize GPU and distributed workloads, reduce unnecessary reruns, and help establish the economics of running reconstruction at scale.
- Work Across the Product Boundary - Partner with Embodied AI product and engineering teams to understand customer requirements and deliver capabilities that fit real training, evaluation, and deployment workflows. Help evolve the representations, tooling, and operational systems that allow reconstructed environments to move reliably into customer applications.
What You'll Bring
- 5+ years of computer vision experience and a bachelor's degree or equivalent,
- Significant experience building and operating production computer-vision, machine-learning, graphics, or data-processing systems.
- Deep hands-on experience with scene understanding: semantic, instance, or panoptic segmentation; open-vocabulary detection and segmentation; or vision-language models applied to perception.
- Experience taking 2D perception into 3D, such as multi-view label fusion, 3D semantic segmentation, instance tracking across views, or scene graph construction over reconstructed geometry.
- Experience delivering reliable computer-vision or 3D systems used by other teams or customers, including the evaluation that proved they worked..
- Practical command of 3D and spatial data, including geometry, camera models, coordinate frames, calibration, and metric scale.
- Strong Python and PyTorch skills, with the ability to work in C++ or other performance-oriented environments when needed.
- Experience with cloud infrastructure, GPU workloads, distributed processing, or large-scale data pipelines.
Nice to Have
- Hands-on experience with 3D reconstruction, photogrammetry, structure from motion, Gaussian splatting, neural rendering, or meshing.
- Experience with SLAM, visual positioning, localization, or large-scale mapping systems.
- Experience with robotics, simulation, or embodied AI applications, in particular how semantic labels are consumed by a policy or a simulator.
- Built evaluation or benchmarking infrastructure for perception or reconstruction systems.
- Experience optimizing large-scale GPU inference workloads for cost, throughput, or latency.
- A publication record or contributions to widely used open-source perception codebases.
- Advanced degree in computer vision, robotics, machine learning, or a related field.
Competencies
- Relentless bias for action. You make decisions and ship as if the company’s success depends on it, because it does. You set the pace, and drive the people around you to rise to it.
- Intellectually honest. You know there’s just one job that underpins every other in a startup: find the truth. You are relentless in your pursuit of it, including when it challenges your own convictions, and you raise concerns when you see them — even when it's uncomfortable, and especially when it's unpopular.
- Pragmatic. You have no patience for not-invented-here Syndrome and analysis paralysis. You find the fastest path to a working capability without one-way doors that undermine scaling.
- Strong systems instincts. You understand that production quality includes reliability, observability, cost, latency, maintainability, and usability—not just algorithmic accuracy.
- Comfort with ambiguity. You can make progress when the problem, data, and requirements are still evolving.
Compensation & Benefits
Base salary range of $229,500 to $255,000 per year. Compensation also includes an annual bonus, equity, and a comprehensive benefits package including medical, dental, and vision coverage, 401(k), and more.
Location & Work Model
This role is based in our San Francisco office, with three days per week in office.
Inclusive Application
We know the strongest candidates don't always tick every box. If you're excited about this role and believe you could do it well, we encourage you to apply even if your experience doesn't match every qualification listed - you may be exactly who we're looking for.
Equal Opportunity
Niantic Spatial is an equal opportunity employer. Individuals seeking employment at Niantic Spatial are considered without regard to race, color, ancestry, national origin, religion, creed, age, gender (including pregnancy, childbirth, breastfeeding or related medical conditions), marital status, physical or mental disability, medical condition, genetic information, military or veteran status, gender identity, gender expression, sexual orientation, or any other protected category under applicable laws. Niantic Spatial will also consider qualified applicants with criminal histories in accordance with applicable laws. Please contact your recruiter if you want to request an accommodation for the job application or interview process.
Candidate Privacy
I understand that by submitting my job application, the information I provide as part of that application will be used in accordance with Niantic Spatial's Privacy Notice for Job Applicants and Candidates https://www.nianticspatial.com/applicant-privacy-notice.
Apply with Beaverhand, the recruiting company that advocates for you and can negotiate your offer