openqareer

ML Researcher — Agentic Reinforcement Learning

Kaishi Partners · Singapore

# ML Researcher — Agentic Reinforcement Learning **Kaishi Partners** · Singapore · `On-site` 🕒 **Статус:** *Опубликовано: вчера* · *Источник: Indeed* --- ### About the Role About the company Our client is an early-stage, venture-backed AI startup building a personal shopping agent that understands users’ preferences, helps them discover relevant products, and supports them in taking action. The team brings together AI agents, personalisation, and commerce to build a service that becomes more useful through ongoing interactions and real customer outcomes. They are hiring in Singapore, with an opportunity for early team members to shape the research agenda and learning systems behind the product. The opportunity You will investigate how an AI agent can make better decisions for a user over time. In commerce, an immediate action is an incomplete measure of success: a purchase may later be returned, preferences can change, and the most useful recommendation may be to buy nothing. Working closely with engineers, you will help define learning objectives, build rigorous evaluations, and develop methods for improving agent behaviour as usable data becomes available. What you’ll do Investigate reward design and credit assignment for delayed outcomes such as satisfaction, returns, and repeat use. Develop approaches to user modelling, memory, and adaptation as preferences and needs change. Explore policies for deciding when an agent should ask, recommend, act, or wait. Build reproducible experiments, meaningful baselines, and ablation studies. Evaluate policy improvements critically, including uncertainty, misleading proxies, and unintended behaviours. Collaborate with engineers on the data and instrumentation needed for research. Translate promising findings into behaviours that can be tested in the product, introducing more sophisticated learning methods when the evidence supports them. What you’ll bring Deep knowledge of reinforcement learning and hands-on research or implementation experience. Strong Python and machine learning engineering skills. Experience designing experiments and evaluating results critically. The ability to translate an open-ended product problem into a tractable research question. An interest in working closely with engineers and learning from real product behaviour. Evidence of research depth through papers, experiments, implementations, or deployed systems. Useful additional experience Sequential decision-making, contextual bandits, or offline reinforcement learning. Recommendation systems, personalisation, or long-term user modelling. Agent evaluation, delayed feedback, or learning from logged interactions. To apply Share your profile and an RL paper, experiment, or implementation you are proud of. Explain your contribution, how you evaluated it, and the limitations of the result. Kaishi Partners (EA No. 16C8316) was set up to meet the recruitment demands of the fast growing South-East Asia technology community - an expansion led by the Singapore government’s ‘Smart Nation’ vision. Our clients span from cutting edge start-ups to global brand names; venture capital to innovation labs; e-commerce platforms to fintech disruptors. If you’re a company anywhere on the spectrum from seed to IPO, or a candidate looking to play your role in making Singapore the next Silicon Valley, then Kaishi are the partner to advise you.

Наблюдалась 2026-09-21, впервые 2026-09-20, источник — Indeed.

Открыть у работодателя