Company: Discovery Health (Discovery Limited)
Industry: Healthcare, Health Insurance, Behavioral Science & Artificial Intelligence
Business Unit: Data Science Lab (DS Lab) / Data Sciences Function
Location: Sandton / Johannesburg, Gauteng
Position Type: Full-Time Permanent / Early Career
Posted Date: 29 July 2026
Closing Date: Open / Rolling Applications
About Discovery Health & The Data Science Lab (DS Lab)
Discovery Health is South Africa’s leading medical scheme administrator, driven by a core purpose to enhance and protect people’s lives through behavioral science, incentives, and data-driven healthcare solutions.
The Data Science Lab (DS Lab) is Discovery’s elite specialist analytics unit. Partnering with global leaders like Google, Quantium Health, and top academic institutions, the DS Lab builds cutting-edge predictive models, causal inference engines, machine learning pipelines, and Agentic AI workflows (such as the Personal Health Pathways programme). These solutions power personalized care and operational decision-making across Discovery South Africa, Vitality UK, Vitality USA, and international markets.
Key Result Areas & Technical Workstreams
Working under the guidance of senior data scientists, your key operational areas will include:
1. Data Analysis, Machine Learning & Causal Inference
- Model Building: Support feature engineering, statistical modeling, machine learning development, and causal inference analysis.
- Measurement Frameworks: Design test-and-learn experiments, impact measurement models, and evaluation frameworks to determine intervention effectiveness.
- Reproducible Code: Maintain clean, reproducible, and well-documented analytical code pipelines.
2. Behavioral Personalization & Experimentation
- Targeting Engine: Develop predictive models using clinical, behavioral, digital, and operational datasets to personalize member interventions.
- A/B Testing: Design and interpret experimental test-and-learn frameworks across health programs.
3. Agentic AI & Advanced Workflows
- GenAI & LLMs: Support the evaluation, build, and deployment of AI-enabled workflows combining Large Language Models (LLMs), Retrieval-Augmented Generation (RAG), structured databases, and business rules.
- AI Reliability: Document risks, failure modes, safety protocols, and commercial viability for AI deployments.
4. Stakeholder Delivery & Business Translation
- Problem Translation: Partner with internal business leaders to translate complex healthcare challenges into quantitative analytics projects.
- Communication: Present analytical findings, recommendations, and model constraints clearly to technical and non-technical stakeholders.
Candidate Profile & Technical Requirements
Candidates must meet the following academic and technical criteria:
- Academic Qualification: A completed Bachelor’s or Honours Degree in a quantitative discipline:
- Data Science / Computer Science / Statistics / Mathematics
- Actuarial Science / Applied Mathematics / Operations Research
- Industrial Engineering (or equivalent technical pathway supported by strong project portfolio).
- Required Technical Stack:
- Python: High proficiency in Python for statistical modeling, machine learning, and data wrangling.
- SQL & Databases: Strong experience querying relational databases and structured data.
- Fundamentals: Solid grounding in statistics, machine learning, experimental design, and model evaluation metrics.
- Advantageous Capabilities:
- Exposure to Causal Inference, Cloud platforms (GCP preferred), Git, or production ML pipelines.
- Exposure to Generative AI, LLMs, RAG, or Agentic AI frameworks.
- Prior experience or academic projects in healthcare, health insurance, behavioral science, or personalization.
- Participation in Kaggle/data science competitions or postgraduate research work.
Career Advice for Discovery Applicants
- Highlight Health & Behavioral Science Projects: Discovery places a huge premium on behavioral incentives (Vitality model) and clinical outcomes. Detail any academic projects or Kaggle entries involving healthcare data, churn modeling, or user behavior prediction.
- Emphasize Python & SQL Proficiency: Ensure your resume clearly details your hands-on experience with Python data packages (Pandas, NumPy, Scikit-Learn, Statsmodels) and SQL queries.
- Mention Generative AI or Causal Inference Exposure: If you have experimented with RAG, LLMs, or causal models (e.g., DoWhy, Uplift Modeling), feature this prominently near the top of your CV to stand out for the DS Lab team.
Sample Interview Preparation Questions
Machine Learning & Causal Inference Questions
- Question (Causal vs. Predictive): “What is the difference between predicting which members will claim high medical expenses versus estimating the causal impact of a wellness intervention on reducing those claims?”
- Question (Model Evaluation): “How do you evaluate a machine learning model when working with highly imbalanced healthcare data (e.g., rare disease or high-cost claim events)?”
Agentic AI & Business Translation
- Question (GenAI Safety): “What risks and failure modes must be documented when deploying Retrieval-Augmented Generation (RAG) or LLM workflows in a healthcare environment?”
How to Apply
Submit your application directly through the official Discovery Careers portal:
- Business Unit: Discovery Health – Data Sciences
- Posted Date: 29 July 2026