Direct-Standardized Rate Difference Between Arms

Description

Computes the difference in a binary outcome rate between the CONTINUE and STOPBASE arms, directly standardized to the pooled sample over a set of stratifying variables, with a nonparametric bootstrap percentile confidence interval. This ports the treatment-pattern secondary analyses of García-Albéniz et al. (SAS cann26/cann27/cann28: surgery, chemotherapy, and radiotherapy receipt among screen-detected cancers), which used PROC STDRATE (method = direct, effect = diff) standardized by age and comorbidity.

Usage

standardized_rate_difference(
  data,
  outcome_col,
  strata_cols,
  arm_col = "arm",
  mult = 1,
  n_boot = 500L,
  conf_level = 0.95,
  seed = NULL
)

Arguments

data A data frame with one row per person (or person-arm), containing the outcome, arm, and stratifying columns.
outcome_col Name of the binary (0/1) outcome column (e.g., receipt of surgery within 12 months).
strata_cols Character vector of stratifying column names (e.g., age group and comorbidity score) defining the standardization strata.
arm_col Name of the trial arm column (“CONTINUE” / “STOPBASE”). Default: “arm”.
mult Rate multiplier; the rates and their difference are scaled by this. Default: 1 (proportions). Use 100 to match the SAS output.
n_boot Number of bootstrap resamples. Default: 500L.
conf_level Width of the percentile confidence interval. Default: 0.95.
seed Optional integer seed; the caller’s RNG state is restored on exit. Default: NULL.

Details

For each arm the stratum-specific outcome rate is computed, then averaged using the pooled-sample stratum distribution as the reference (direct standardization). The standardized rate difference is rate(CONTINUE) - rate(STOPBASE). Standardization is restricted to strata present in both arms (common support); the reference weights are rescaled to sum to one over those strata. Multiply by mult to report a rate per mult people (the SAS code used mult = 100).

The confidence interval is the central conf_level percentile range of the standardized rate difference across n_boot bootstrap resamples of the rows (reference weights are recomputed on each resample). The caller’s RNG state is restored on exit.

Value

A named list with rate_continue and rate_stopbase (the standardized arm rates), diff (their difference), conf_low, conf_high, n_boot, and conf_level. All rates are scaled by mult.

References

García-Albéniz X, Uno H, Bhatt DL, McArdle PH, Joffe MM, Hernán MA. Continuation of Annual Screening Mammography and Breast Cancer Mortality in Women Older Than 70 Years: A Prospective Observational Study. Ann Intern Med. 2020;172(6):381-389. doi:10.7326/M18-1199

Examples

Code
library("ettbc")

set.seed(1)
n <- 400
dat <- data.frame(
  arm = rep(c("CONTINUE", "STOPBASE"), each = n / 2),
  age_group = sample(c("70-74", "75-84"), n, replace = TRUE),
  surgery = rbinom(n, 1, 0.5)
)
standardized_rate_difference(
  dat, "surgery", strata_cols = "age_group", n_boot = 200, seed = 1
)
$rate_continue
[1] 0.5147794

$rate_stopbase
[1] 0.5145522

$diff
[1] 0.0002272065

$conf_low
[1] -0.09985649

$conf_high
[1] 0.09613983

$n_boot
[1] 200

$conf_level
[1] 0.95