Modeling and analyzing respondent-driven sampling as a counting process

DOI10.1111/BIOM.12678MaRDI QIDQ4556702zbMATH OpenOpenAlexWikidataFDO

Authors Yakir Berchenko, Jonathan D. Rosenblatt, Simon D. W. Frost

Publication date 16 November 2018

Published in Biometrics (Search for Journal in Brave)

Full work available at URL https://arxiv.org/abs/1304.3505

counting process HIV respondent driven sampling hidden populations

Applications of statistics to biology and medical sciences; meta analysis (62P10) Sampling theory, sample surveys (62D05)

Abstract: Respondent-driven sampling (RDS) is an approach to sampling design and analysis which utilizes the networks of social relationships that connect members of the target population, using chain-referral methods to facilitate sampling. RDS typically leads to biased sampling, favoring participants with many acquaintances. Naive estimates, such as the sample average, which are uncorrected for the sampling bias, will themselves be biased. To compensate for this bias, current methodology suggests inverse-degree weighting, where the "degree" is the number of acquaintances. This stems from the fundamental RDS assumption that the probability of sampling an individual is proportional to their degree. Since this assumption is tenuous at best, we propose to harness the additional information encapsulated in the time of recruitment, into a model-based inference framework for RDS. This information is typically collected by researchers, but ignored. We adapt methods developed for inference in epidemic processes to estimate the population size, degree counts and frequencies. While providing valuable information in themselves, these quantities ultimately serve to debias other estimators, such a disease's prevalence. A fundamental advantage of our approach is that, being model-based, it makes all assumptions of the data-generating process explicit. This enables verification of the assumptions, maximum likelihood estimation, extension with covariates, and model selection. We develop asymptotic theory, proving consistency and asymptotic normality properties. We further compare these estimators to the standard inverse-degree weighting through simulations, and using real-world data. In both cases we find our estimators to outperform current methods. The likelihood problem in the model we present is convex, and thus efficiently solvable. We implement these estimators in an R package, chords, available on CRAN.

Recommendations

Cited in

(15)

Describes a project that uses

Uses Software

This page was built for publication: Modeling and analyzing respondent-driven sampling as a counting process

Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q4556702)