Hiring.Camp

Research Intern – Reinforcement Learning for Large Foundation Models

Tencent

·

1 week ago

Location
Singapore-CapitaSky
Workplace
Onsite
Type
Internship
Department
Education
Seniority
Internship
Education
PhD
Source
Workday

Description

Business Unit

Technology Engineering Group (TEG) is responsible for supporting the company and its business groups on technology and operational platforms, as well as the construction and operation of R&D management and data centers, TEG provides users with a full range of customer services. As the operator of the largest networking, devices, and data center in Asia,TEG also leads the Tencent Technology Committee in strengthening infrastructure R&D through internal and distributed open source collaboration, constructing new platforms and supporting business innovation.

What the Role Entails

Research directions include but are not limited to:
- RL Algorithms for Reasoning Models:Design robust RL training recipes (PPO / GRPO / GSPO variants) for large-scale reasoning models. Tackle training instability, reward hacking, and policy collapse in long-horizon and async settings. Explore how to bridge the gap between RL post-training and genuine reasoning capability improvement.
- RL for Autonomous Agents: Build RL pipelines for long-horizon terminal agents and tool-use agents. Investigate credit assignment, exploration strategies, and self-evolving agent behaviors in complex interactive environments.
-Reward Modeling & Optimization: Develop reward signals and regularization techniques that go beyond outcome-based rewards. Explore token-level reward shaping, entropy-based regularization, and learned reward models that generalize across tasks.
 

Who We Look For

- Enrolled in a PhD or Master's program in computer science, machine learning or a related field. 
- Solid understanding of RL fundamentals (policy gradients, PPO, GRPO, etc.) and hands-on experience applying them to LLM training.
- Strong programming skills in Python and PyTorch; experience training models on multi-GPU setups, comfortable debugging training instability at scale.
- Ability to independently read, critique, and build on recent research papers.
- Published or submitted first-author papers at top ML/NLP venues (NeurIPS, ICML, ICLR, ACL, AAAI, EMNLP, etc.), or demonstrated equivalent research maturity through preprints and technical reports.
- Familiarity with LLM post-training (RLHF, DPO, GRPO), model merging, or agent frameworks (ReAct, tool-use) is a strong plus.
- Experience with large-scale distributed training (DeepSpeed, FSDP, Megatron) and open-source contributions are welcome.

Equal Employment Opportunity at Tencent

As an equal opportunity employer, we firmly believe that diverse voices fuel our innovation and allow us to better serve our users and the community. We foster an environment where every employee of Tencent feels supported and inspired to achieve individual and common goals.

Skills

PythonMachine LearningNLPPyTorch

Similar Jobs

30

Research Intern

Nshs · NSO 1001 University Place Evanston, United States of America · Onsite

4 days ago

Research Intern

bannerhealth · BUMC Phoenix (1111 E McDowell Rd), United States of America

2 weeks ago

Research Intern

AHRI · AHRI - Durban, South Africa

2 weeks ago

Research Intern

Integritymarketing · DFT - Minneapolis, MN, United States of America

2 weeks ago

Intern, Research

Fred Hutch · Seattle, WA, US

2 weeks ago

Research Intern

HSS - Hospital for Special Surgery · HSS 777, United States of America · Hybrid

3 weeks ago

Research Intern

Smithfieldfoods · SPG R&D Warsaw NC - HP SPG, United States of America · Onsite

4 weeks ago

Research Intern

Smithfieldfoods · SPG R&D Warsaw NC - HP SPG, United States of America · Onsite

4 weeks ago

Research Intern

HSS - Hospital for Special Surgery · HSS 777, United States of America · Hybrid

1 month ago

Research Intern

Iterativehealth · Tacoma, WA +1 · Onsite

1 month ago

Research Intern

Altimate · Bengaluru · remote

1 month ago

Research Intern

Nationwidechildrens · 575 Children's Crossroad RB4, United States of America

1 month ago

Research Intern

Rgare · China, Shanghai · Onsite

1 month ago

Research Intern

Nshs · EVH Burch Building Evanston Hospital, United States of America · Remote, Onsite

1 month ago

Intern-Research

NemoursCareerSite · Wilmington, DE, United States, US · Onsite

1 month ago

Research Intern

New York Blood Center · Rye, NY, US

1 month ago

Research Intern

Fred Hutch · Seattle, WA, US

1 month ago

Intern, Research

Fred Hutch · Seattle, WA, US

1 month ago

Intern, Research

Fred Hutch · Seattle, WA, US

1 month ago

Research Intern

Usuro · US NY Water Street, United States of America

1 month ago

Research Intern

Usuro · US NY Water Street, United States of America

1 month ago

Research Intern

Fred Hutch · Seattle, WA, US

2 months ago

Intern, Research

Fred Hutch · Seattle, WA, US

2 months ago

Research Intern

Cedars-Sinai Medical Center · Los Angeles, CA, United States, US

2 months ago

Research Intern

Nokia · India · Onsite

2 months ago

Research Intern

Gia · Carlsbad Headquarters, United States of America · Onsite

2 months ago

Research Intern

Nationwidechildrens · 611 E Livingston Avenue, United States of America

2 months ago

Intern-Research

NemoursCareerSite · Wilmington, DE, United States, US · Remote

2 months ago

Research Intern

Fred Hutch · Seattle, WA, US

2 months ago

Research Intern

Psu · Penn State University Park, United States of America · Remote, Hybrid, Onsite

2 months ago
Research Intern – Reinforcement Learning for Large Foundation Models at Tencent | Hiring.Camp