西门子Expert Research Scientist on Autonomous Polyfunctional Robots (Foundation Models)
社招全职5-10年研发地点:苏州状态:招聘♡ 收藏
工作描述
职位概述
你将负责通用具身基础模型的全周期评测、领域适配与性能优化,包括 VLA 模型和端到端策略模型,并推动它们从训练完成走向真实世界部署。你所优化的基础模型不仅将支撑机器人的认知、感知、推理与规划能力,也将面向全身控制(WBC),帮助机器人实现协调、稳定且符合物理约束的运动。在这个岗位上,你将围绕仿真与实体机器人构建统一评测框架,推进后训练与微调研究,并主导架构 Physical AI Harness:一套标准化的真实机器人评测基础设施、自动化测试框架与安全护栏系统。你将直面 Sim2Real 迁移、跨本体适配、长时程任务鲁棒性、WBC 可靠性以及物理世界安全部署等关键挑战,直接提升机器人在真实场景中的任务成功率、泛化能力与可靠性。
工作职责
• 评测框架与基准体系:面向仿真和真实机器人环境,设计并工程化具身模型的综合评测框架。定义多维度指标,例如任务成功率、轨迹平滑性、碰撞率、物理一致性以及长时程推理准确率。建设自动化基准测试平台和标准化评测套件。
• 后训练与领域适配:主导通用 VLA 与策略模型的领域适配,开发监督微调、强化学习微调、偏好对齐和课程学习策略,将通用模型能力高效迁移到具体任务、机器人本体以及全身控制场景中。
• Physical AI Harness 架构与运营:设计并构建 Physical AI Harness,形成标准化真实机器人测试基础设施,覆盖多场景评测测试台、自动化任务编排、执行框架以及快速跨本体适配接口。建立物理世界评测所需的安全约束与回退机制,包括异常检测、碰撞避免、力限保护和紧急停止策略。构建真实机器人 A/B 测试和持续评测流水线,支持模型版本的回归测试、性能漂移监测和部署决策。
• 跨本体与 Sim2Real 优化:开发跨本体微调方法,解决不同机器人之间动作空间异构和观测空间不一致的问题。诊断并缩小 Sim2Real 差距,利用 Physic…登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
强化学习+
https://cloud.google.com/discover/what-is-reinforcement-learning?hl=en
Reinforcement learning (RL) is a type of machine learning where an "agent" learns optimal behavior through interaction with its environment.
https://huggingface.co/learn/deep-rl-course/unit0/introduction
This course will teach you about Deep Reinforcement Learning from beginner to expert. It’s completely free and open-source!
https://www.kaggle.com/learn/intro-to-game-ai-and-reinforcement-learning
Build your own video game bots, using classic and cutting-edge algorithms.
学历+
深度学习+
https://d2l.ai/
Interactive deep learning book with code, math, and discussions.
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
Transformer+
https://huggingface.co/learn/llm-course/en/chapter1/4
Breaking down how Large Language Models work, visualizing how data flows through.
https://poloclub.github.io/transformer-explainer/
An interactive visualization tool showing you how transformer models work in large language models (LLM) like GPT.
https://www.youtube.com/watch?v=wjZofJX0v4M
Breaking down how Large Language Models work, visualizing how data flows through.
SFT+
https://cameronrwolfe.substack.com/p/understanding-and-using-supervised
Understanding how SFT works from the idea to a working implementation...
RLHF+
[英文] What is RLHF?
https://aws.amazon.com/what-is/reinforcement-learning-from-human-feedback/
Reinforcement learning from human feedback (RLHF) is a machine learning (ML) technique that uses human feedback to optimize ML models to self-learn more efficiently.
https://www.ibm.com/think/topics/rlhf
Reinforcement learning from human feedback (RLHF) is a machine learning technique in which a “reward model” is trained with direct human feedback, then used to optimize the performance of an artificial intelligence agent through reinforcement learning.
TypeScript+
https://www.youtube.com/watch?v=JHEB7RhJG1Y
Master TypeScript from basics to advanced concepts through hands-on tutorials covering type annotations, generics, data fetching, Zod library, and more, with practical challenges for effective real-world application.
还有更多 •••
相关职位
社招5-10年研发
职位概述 作为连接数据基础设施与算法创新的关键桥梁,你将端到端负责具身智能数据栈的建设。你将设计高吞吐、高可靠的多模态数据管线和 Physical AI Harness,并定义数据与算法策略,包括数据
更新于 2026-09-30苏州
社招5-10年研发
The Opportunity: Expert Research Scientist on Polyfunctional Robot Perception & AI Are you a tra
更新于 2026-07-03上海|苏州
社招5-10年研发
The Opportunity: Expert Research Scientist on Polyfunctional Robot Mechatronics System Are you an ex
更新于 2026-07-03上海|苏州