阿里巴巴集团安全部-强化学习/Agent算法工程师/专家-行为风控方向
社招全职3年以上地点:北京状态:招聘
工作描述
任职要求 1、硕士研究生及以上学历,计算机、人工智能、软件、信息安全、统计和数学专业优先; 2、3年以上大模型/强化学习相关研发经验,深刻理解RLHF/Agent训练经验; 3、具备业务风控领域(作弊、欺诈、账号安全、恶意行为等方向)的实战经验,对风险数据(日志、行为序列、用户画像、图数据)有敏锐的洞察力和处理经验者优先; 4、扎实的编程基础,对数据结构、算法设计基础有深度了解,熟练掌握Pyth…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
学历+
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
强化学习+
https://cloud.google.com/discover/what-is-reinforcement-learning?hl=en
Reinforcement learning (RL) is a type of machine learning where an "agent" learns optimal behavior through interaction with its environment.
https://huggingface.co/learn/deep-rl-course/unit0/introduction
This course will teach you about Deep Reinforcement Learning from beginner to expert. It’s completely free and open-source!
https://www.kaggle.com/learn/intro-to-game-ai-and-reinforcement-learning
Build your own video game bots, using classic and cutting-edge algorithms.
RLHF+
[英文] What is RLHF?
https://aws.amazon.com/what-is/reinforcement-learning-from-human-feedback/
Reinforcement learning from human feedback (RLHF) is a machine learning (ML) technique that uses human feedback to optimize ML models to self-learn more efficiently.
https://www.ibm.com/think/topics/rlhf
Reinforcement learning from human feedback (RLHF) is a machine learning technique in which a “reward model” is trained with direct human feedback, then used to optimize the performance of an artificial intelligence agent through reinforcement learning.
还有更多 •••
相关职位
社招3年以上技术类-安全
1.具有攻防经验,熟悉常见漏洞原理,善于漏洞的挖掘、利用 2.熟悉RASP相关技术原理,有 RASP 建设经验者优先 3.具备良好的技术洞察力和创新能力,能够快速识别问题、解决问题 4.具备良好的风险
更新于 2026-06-18北京|杭州
社招2年以上技术类-算法
1. 计算机、信息安全、人工智能或相关专业硕士及以上学历; 2. 具备扎实的大模型基础知识和训练经验,熟悉大模型的模型架构和改进方案,熟悉预训练、SFT、GRPO、PPO、Agentic RL 等主流
更新于 2026-07-03北京
社招3年以上技术类-算法
1、计算机、人工智能或相关专业硕士及以上学历 2、3年以上大模型/智能体算法经验,精通大模型和智能体的基本原理与训练方法 3、具备优秀的结构化思维和信息提炼能力,善于沟通和跨团队协作 4、在大模型、A
更新于 2026-06-17北京|杭州