Fit the Screening-Propensity Model

Description

Fits a pooled logistic regression for the probability of receiving a screening mammogram at each eligible participant-month, and returns the fitted model together with the predicted probabilities merged onto the input data. These predictions are the pred_prob_col consumed by compute_ipw_weights().

Usage

fit_screening_propensity(
  long_data,
  covariates = character(0),
  scrmammo_col = "scrmammo",
  tslm_lag_col = "tslm_lag",
  month2_col = "month2",
  id_col = "id",
  rcs_knots = c(13, 16, 25, 27),
  min_tslm_lag = 11L,
  pred_col = "p_scrmammo"
)

Arguments

long_data A data frame in long format augmented by augment_long_covariates(). Must contain the scrmammo, tslm_lag, and month2 columns (or as specified via the *_col arguments), plus every column named in covariates.
covariates Character vector of additional covariate column names to include in the model. Default: none. Supply the baseline and time-varying adjustment covariates here when analyzing real cohort data.
scrmammo_col Name of the binary screening-mammogram outcome column. Default: “scrmammo”.
tslm_lag_col Name of the lagged time-since-last-mammogram column. Default: “tslm_lag”.
month2_col Name of the 0-indexed month-from-entry column. Default: “month2”.
id_col Name of the participant ID column. Default: “id”.
rcs_knots Numeric vector of restricted-cubic-spline knots for tslm_lag. Default: c(13, 16, 25, 27) (the SAS tslm_lagII knots).
min_tslm_lag Minimum tslm_lag for a row to enter the model fit and receive a prediction. Default: 11L.
pred_col Name of the predicted-probability column to add to the returned data. Default: “p_scrmammo”.

Details

This ports the SAS cann17b denominator (%cann17b_all_model) propensity model. The outcome is the screening-mammogram indicator scrmammo. The linear predictor combines:

  • the time since the last mammogram, tslm_lag, as both a linear term and a restricted-cubic-spline basis (knots at rcs_knots, using the same Harrell parameterization as predict_survival_ipw());

  • month from entry as month2 and month2^2;

  • any additional baseline or time-varying covariates named in covariates.

The model is fit only on the decision window, the rows with tslm_lag >= min_tslm_lag, matching the SAS where tslm_lag >= 11. Because the propensity model is arm-independent (the SAS model is fit on the uncloned person-time), the fitting sample is deduplicated to one row per participant-month before fitting, so a participant is not counted twice for appearing in both arms.

Predicted probabilities are returned for every row in the decision window; rows outside it (including trial entry, where tslm_lag is NA) receive NA. compute_ipw_weights() treats those rows as having no predicted screening, consistent with the SAS weights program.

Value

A list with two elements:

  • model: The fitted glm object (binomial family, logit link).

  • data: long_data with the pred_col column added (NA outside the tslm_lag >= min_tslm_lag decision window).

References

García-Albéniz X, Uno H, Bhatt DL, McArdle PH, Joffe MM, Hernán MA. Continuation of Annual Screening Mammography and Breast Cancer Mortality in Women Older Than 70 Years: A Prospective Observational Study. Ann Intern Med. 2020;172(6):381-389. doi:10.7326/M18-1199

See Also

augment_long_covariates() for the preceding step and compute_ipw_weights() for the step that consumes pred_col.

Examples

Code
library("ettbc")

cloned <- clone_censor(cohort, screening_mammograms, diagnostic_mammograms)
long_data <- expand_to_long(cloned)
long_data <- augment_long_covariates(
  long_data,
  screening_mammograms,
  diagnostic_mammograms
)
fit <- fit_screening_propensity(long_data)
weighted <- compute_ipw_weights(fit$data, pred_prob_col = "p_scrmammo")
head(weighted[, c("id", "arm", "month2", "w", "wp99")])
  id      arm month2 w wp99
1  1 STOPBASE      0 1    1
2  1 STOPBASE      1 1    1
3  1 STOPBASE      2 1    1
4  1 STOPBASE      3 1    1
5  1 STOPBASE      4 1    1
6  1 STOPBASE      5 1    1