---
title: "Calculating Health Inequality in Complex Surveys"
author: "Sujon Mia - StatAid Research Lab"
output: rmarkdown::html_vignette
vignette: >
  %\VignetteIndexEntry{Calculating Health Inequality in Complex Surveys}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r, include = FALSE}
is_srvyr_ready <- requireNamespace("srvyr", quietly = TRUE)
knitr::opts_chunk$set(
  collapse = TRUE,
  comment = "#>",
  eval = is_srvyr_ready
)
```

## Introduction

Analyzing health inequalities—such as the concentration of childhood malnutrition among the poorest wealth quintiles—is a critical component of global health research. However, calculating standardized metrics like the **Concentration Index** using data from complex, multi-stage stratified cluster surveys (such as DHS or MICS) is methodologically challenging. 

The `SurveyNCD` package, developed by the **StatAid Research Lab**, provides a streamlined, mathematically rigorous pipeline to automate these calculations.

This vignette demonstrates how to process raw anthropometric data and calculate survey-weighted health inequality metrics in just a few lines of code.

## Setup and Package Loading

First, load `SurveyNCD` along with the standard `tidyverse` and `srvyr` packages for data manipulation and survey design handling.

```{r setup-packages, message=FALSE, warning=FALSE}
library(SurveyNCD)
library(dplyr)
library(srvyr)
```

## Phase 1: Cleaning Anthropometric Data

Raw survey datasets often store Z-scores as integers (e.g., `-254` instead of `-2.54`) and use specific numerical flags (like `9999`) to indicate missing or biologically implausible data. Failing to account for these flags will severely distort your means and indices.

The `who_anthro_score()` function automatically cleans these values and categorizes them according to strict WHO child growth standards.

```{r anthro-clean}
# Simulating raw survey data with a missing flag (9999)
raw_survey_data <- tibble(
  cluster_id = c(1, 1, 2, 2),
  strata = c(1, 1, 2, 2),
  sample_weight = c(1.2, 0.8, 1.1, 0.9),
  wealth_score = c(-1.5, -0.2, 0.5, 1.8),
  raw_hw70 = c(-310, -254, 50, 9999) # Raw Height-for-Age
)

# Apply the SurveyNCD cleaning function
cleaned_data <- raw_survey_data %>%
  mutate(
    stunting_category = who_anthro_score(raw_hw70, indicator = "stunting"),
    is_stunted = case_when(
      stunting_category %in% c("Severe stunting", "Moderate stunting") ~ 1,
      stunting_category == "Normal stunting" ~ 0,
      TRUE ~ NA_real_
    )
  )

cleaned_data %>% select(raw_hw70, stunting_category, is_stunted)
```

## Phase 2: Specifying the Complex Survey Design

Before calculating inequality, we must declare the complex survey design.

```{r survey-design}
svy_design <- cleaned_data %>%
  as_survey_design(
    ids = cluster_id,
    strata = strata,
    weights = sample_weight
  )
```

## Phase 3: Calculating the Concentration Index

The `survey_concentration_index()` function handles the internal complexities of calculating weighted fractional ranks and covariances across the survey design.

```{r calc-ci}
stunting_inequality <- svy_design %>%
  survey_concentration_index(
    outcome = is_stunted,
    wealth = wealth_score
  )

print(stunting_inequality)
```

### Interpretation
The output provides both the weighted prevalence of the outcome and the precise `Concentration_Index`. A negative index confirms a pro-poor inequality in child stunting.
