Senior AI Evaluation Engineer #AIDA
Date: 21 Sept 2026
Location: Singapore, Singapore
Company: Singtel Group
Powering the Future with AIDA
To lead the next phase of our AI evolution, we’ve launched a new business unit AIDA – Artificial Intelligence & Data Analytics – a strategic engine driving our transformation designed to scale our AI ambitions with precision and purpose. This marks a pivotal shift in how we operate, innovate, and serve to embed intelligence into every layer of our business.
At Singtel, this is more than a technology upgrade. It’s a strategic transformation that redefines how value is created across the enterprise core—augmenting human capabilities and unlocking entirely new potential. It is a transformation journey by aligning people, platforms, and processes under one cohesive strategy. Our mission is to build AI literacy and foster a culture where intelligence empowers people.
We welcome you to join us on a transformational journey that’s reshaping the telecommunications industry — and redefining what’s possible with AI at its core. Grow with us in a workplace that champions innovation, embraces agility, and puts human potential at the heart of everything we do.
Summary of the role:
Serve as the senior technical evaluator within the team, responsible for the design, calibration, and adjudication of evaluation suites and gate thresholds across agent archetypes. Lead gate reviews under the Team Lead’s authority, own the regression-pack methodology for AI/ML model changes, and act as the technical custodian of evaluation quality and drift hygiene.
How You will Make An Impact:
- Design and maintain offline evaluation suites (golden sets, regression packs, adversarial/safety probes) and the continuous-evaluation scoring pipeline across archetypes.
- Conduct operability gate reviews: assess evidence packs, reproduce evaluation results, and recommend go/no-go with documented findings.
- Own the model-update regression pack methodology and adjudicate regression runs against archived baselines with AIML Operations team.
- Calibrate gate thresholds against production reality and maintain evaluation drift hygiene (goldenset rotation, hold-out sets, judge calibration).
- Produce the monthly quality report per agent: eval trends, failure-mode taxonomy, and defect clusters with reproduction traces.
- Mentor members of the team when needed and review their work for quality and consistency.
Skills for Success:
- Bachelor’s or Master’s degree in Computer Science or a related field
- 6+ years in ML/data/software with strong evaluation or quality focus
- Hands-on experience evaluating LLM or ML systems
- LLM/agent evaluation design and statistical rigour
- Python and evaluation tooling (promptfoo, DeepEval, or custom harnesses)
- Data analysis and metric interpretation
- Tracing/observability tooling
- Adversarial testing / red teaming Non-Technical / Soft Skills
- Analytical rigour and attention to detail
- Clear technical writing for gate findings
- Ability to influence build teams on quality
Good to have:
- Experience with agentic systems or RAG pipelines
- Telco AI domain
- Adversarial testing / red teaming
- Management presentation
- Responsible-AI and safety evaluation practices
- Understanding of telco customer intents and journeys
Are you ready to say hello to BIG Possibilities?
Join Singtel to shape what's next and accelerate your career through meaningful work, continuous learning, and real impact.