ettbc (development version)

  • Updated the reusable GitHub Actions workflows to call Morrison-Lab/gha instead of d-morrison/gha, following that repository’s move. Actions does not follow repository-rename redirects for uses:, so the calls were failing to resolve before any job started.

  • Added apply_eligibility_criteria(), the cohort eligibility/enrollment logic from García-Albéniz et al. (SAS c01_eligibility.sas, item 3). Given a demographics table, a monthly enrollment table, and screening mammogram events, it selects participants alive at the age threshold, with a qualifying mammogram near that age (which becomes the derived study entry month), twelve consecutive months of fee-for-service Medicare enrollment ending at entry, and entitlement by age rather than disability or end-stage renal disease. Complements gagne_weights() / comorbidity_score() (#27) as the other half of cohort construction (#31, #12).

  • fit_weighted_logistic() now muffles the expected non-integer #successes in a binomial glm! warning that non-integer IPW weights trigger. The fit is unchanged and other warnings (non-convergence, separation) still surface; this just stops bootstrap_ci() from emitting one such warning per resample.

  • Added fit_weighted_logistic(): the shared weighted pooled-logistic fitting primitive behind fit_outcome_hr() and predict_survival_ipw(), which both now call it instead of repeating the weighted-GLM block. Exported so sibling packages can reuse the same fit; first step of the cross-package outcome-model consolidation (#25). No change to existing results (#12).

  • Added gagne_weights() and comorbidity_score(), the Gagne combined comorbidity score used by García-Albéniz et al. for cohort eligibility and adjustment (item 3). gagne_weights() returns the 20 published per-condition weights from the official ICD-9-CM / ICD-10-CM combined-comorbidity-score program – including the two negative weights (HIV/AIDS and hypertension). comorbidity_score() applies a named weight vector to per-person 0/1 condition indicators, defaulting to the Gagne weights. Mapping ICD codes to the condition flags is left to the user (e.g. the {comorbidity} package) (#12).

  • bootstrap_ci() now seeds its resampling with withr::with_seed() instead of a hand-rolled .Random.seed save/restore, and its bootstrap loop was factored into a helper. Behavior and results are unchanged (#12).

  • Added standardized_rate_difference(), the treatment-pattern secondary analyses from García-Albéniz et al. (SAS cann26/cann27/cann28: surgery, chemotherapy, and radiotherapy receipt among screen-detected cancers). It computes a direct-standardized difference in a binary outcome rate between the CONTINUE and STOPBASE arms, standardized to the pooled sample over user-specified strata (age, comorbidity) and restricted to the common-support strata, with a bootstrap percentile confidence interval and RNG-state restoration, matching the SAS PROC STDRATE (method = direct, effect = diff) approach (#12).

  • Added deterministic_bias_analysis() and probabilistic_bias_analysis(), the unmeasured-confounding sensitivity analyses from García-Albéniz et al. (supplementary analysis). For a single dichotomous, time-fixed confounder with a given prevalence in each arm and an additive outcome effect, the deterministic version adjusts an observed arm risk difference by (prev_continue - prev_stopbase) * confounder_effect (vectorized over a grid of assumptions); the probabilistic version draws the bias parameters from uniform priors and the observed risk difference from its sampling distribution to return a Monte Carlo simulation interval. The caller’s RNG state is restored on exit (#12).

  • Exported compute_rcs_basis(), the Harrell restricted-cubic-spline basis (SAS %RCSPLINE parameterization) used internally by the outcome and propensity models, so other packages can reuse the same parameterization instead of writing it again. Behavior is unchanged; the function moved to its own file and gained a public man page, examples, and tests (#12).

  • Added negative_control_analysis() and negative-control outcome support, the falsification check from García-Albéniz et al. (death from cancer of the corpus uteri). expand_to_long() gains an optional nc_died_col argument that builds a cause-specific nc_dead_t1 outcome the same way as bc_dead_t1 (the shared logic is factored into one helper); simulate_screening_cohort(negative_control = TRUE) adds an nc_death indicator to the simulated cohort. negative_control_analysis() runs fit_outcome_hr() on the negative-control outcome and reports null_consistent, whether the arm effect’s confidence interval covers the null. Continued screening should have no effect on a death unrelated to the breast; a clearly non-null result would flag residual bias. The default behavior of expand_to_long() and simulate_screening_cohort() is unchanged (#12).

  • Added simulate_screening_cohort(): an exported generator that simulates a synthetic cohort of arbitrary size and returns the three linked data frames (cohort, screening_mammograms, diagnostic_mammograms) the pipeline needs. It wraps the internal generators behind a single seeded random-number stream, so simulate_screening_cohort(100, 108, seed = 2020) reproduces the shipped example datasets exactly; the seed is applied with withr::with_seed(), so the caller’s RNG stream is left untouched. The “Using ettbc” article uses it to demonstrate the full clone_censor() -> compute_ipw_weights() -> fit_outcome_hr() -> predict_survival_ipw() -> bootstrap_ci() pipeline end to end on a larger simulated cohort (#12).

  • Added augment_long_covariates(): builds the time-varying screening covariates the weight and propensity steps need from the long-format data and the mammogram events, porting the SAS cann17b augmentation. It adds scrmammo, dxmammo, anymammo, tslm, tslm_lag, and monthBC, forcing a screen at entry, resetting the time-since-last-mammogram clock at any mammogram, not counting a screen in the month after a breast-cancer diagnosis, and reclassifying a screen within dx_reclass_months of the previous mammogram as diagnostic (#12).

  • Added fit_screening_propensity(): fits the pooled logistic screening-propensity model (SAS cann17b denominator) and returns the predicted p_scrmammo that compute_ipw_weights() consumes. The linear predictor uses tslm_lag (linear plus a restricted-cubic-spline basis), month2 and month2^2, and any user-supplied baseline/time-varying covariates; the model is fit on the tslm_lag >= min_tslm_lag decision window, deduplicated to one row per participant-month. Together with augment_long_covariates() this lets clone_censor() output drive compute_ipw_weights() end to end without hand-set columns (#12).

  • Removed leftover package-template scaffolding: the example_function() function (and its test and man page), the quarto_vignette.qmd and quarto_article.qmd template-demo vignettes, and the generic CHECKLIST.md and USAGE.md template setup guides. Rewrote README (.Rmd/.md) and the “Getting Started” vignette to describe {ettbc} and demonstrate the clone_censor() -> expand_to_long() pipeline on the synthetic example data, and replaced the placeholder inst/extdata/README.md.

  • Added predict_survival_unadjusted(), predict_survival_baseline_adjusted(), and predict_survival_ipw(): fit pooled logistic regression models with restricted cubic spline time terms and arm-by-time interactions, then apply g-computation to produce marginal survival curves for each trial arm. Time is modeled with a full-rank Harrell restricted-cubic-spline basis (a linear month3 term plus length(rcs_knots) - 2 nonlinear terms) following the SAS %RCSPLINE macro, rather than splines::ns(), which previously aliased with the linear term and left the design rank-deficient. The marginal survival curves are unchanged; the spline terms now vanish at month 0, so the fit_outcome_hr() arm odds ratio is the well-defined contrast at baseline. All three functions apply the max_month filter before model fitting, require both arms to be present (before and after filtering), and validate weight_col when supplied (#1).

  • Added compute_ipw_weights(): computes stabilized cumulative IPW weights for each participant-arm-month, truncated at the 99th percentile computed separately within each arm. CONTINUE-arm weight logic matches the SAS cann17b implementation: updates occur at every month in the tslm_lag 11–13 compliance window (not only when scrmammo == 1), using conditional uniform probabilities (1/3, 1/2, 1 at months 11, 12, 13 respectively), and stop after a breast-cancer diagnosis (#1).

  • Added fit_outcome_hr(): fits an IPW-weighted pooled logistic regression and returns the odds ratio (hazard ratio approximation) for the STOPBASE arm with a 95% Wald CI. Uses cluster-robust variance via sandwich::vcovCL() when sandwich is available (new cluster_id_col argument). Validates both weight_col and cluster_id_col before fitting. Added sandwich to Suggests (#1).

  • Added bootstrap_ci(): nonparametric bootstrap for 95% percentile confidence intervals on the IPW-estimated survival difference. Failed iterations are counted; a warning is issued when the failure rate exceeds fail_threshold (default 10%). The caller’s RNG state is saved and restored on exit. Validates that n_boot is a positive integer (#1).

  • Added false_positives(): computes false positive rates for histological evaluations, stratified by trial arm and screening round. Filters evaluations to within each arm’s observed follow-up and deduplicates repeat evaluations within window_months per participant-arm (#1).

  • Added extract_screening_mammograms(), extract_any_mammograms(), and extract_diagnostic_mammograms(): template functions for extracting mammogram events from Medicare claims data by HCPCS code (#1).

  • Added cli to Imports in DESCRIPTION (the restricted cubic spline basis is computed directly, so splines is no longer a dependency).

  • Added clone_censor(): implements the clone-censor step of the target trial emulation methodology. Creates two clones per participant (STOPBASE and CONTINUE arms) and applies the corresponding censoring rules (#1).

  • Added expand_to_long(): converts cloned data to one row per participant-arm-month for use in discrete-time survival analysis (#1).

  • Added example synthetic datasets: cohort, screening_mammograms, and diagnostic_mammograms (#1).

  • Added vignette article “Using ettbc: Emulating a Target Trial for Breast Cancer Screening” (#1).

  • Updated DESCRIPTION title and description to reflect the package purpose.

  • Migrated the @claude, Claude review, and NEWS.md changelog-check GitHub Actions workflows to the reusable workflows in d-morrison/gha.

  • Internal refactor: decomposed the per-participant clone-censor logic in clone_censor() into smaller helper functions, factored the cumulative mortality computation in the “Using ettbc” vignette into a reusable helper, and split the example-data generation into focused simulation functions now living in R/. Helper functions were reorganized to one function per file, anonymous functions replaced with named helpers, and nested calls rewritten as pipes. No user-facing behavior changes.

ettbc 0.0.0.9000

  • Initial development version