1. Overview

Sampling “hard-to-reach” or “rarely heard”, populations has always been a challenge to researchers, and this has had negative consequences for those populations, as policy making and legislation are less efficient when there is little high-quality data. Respondent-driven sampling (RDS) combines “snowball sampling” (where existing participants recruit future participants from their network of friends or acquaintances) with a dedicated statistical estimator that adjusts for certain biases associated with snowball sampling. Almost everyone now owns a smartphone, and this has allowed for the smartphone-based respondent-driven sampling (SB-RDS) method to be developed. The SB-RDS method involves participants receiving a link to the survey by text and completing the survey on their smartphone. Upon completion, their reward is sent by text, with an invitation code the participant can share with prospective participants in their social circle.

The aim of the project was to establish, via a series of real-life tests, whether SB-RDS could become a useful tool for the Office for National Statistics (ONS). The study ran from 26 February 2024 until 31 March 2025.

The target populations for the three tests were:

  • residents of York City Council, aged 16 years and over
  • people identifying as Gypsy, Roma or Traveller, living in England and aged 16 years and over
  • People living on a boat for at least three months in a year, living in England and aged 16 years and over

The results from the York test were compared with census values and this showed that confidence intervals of the RDS estimates contained the census values for sex but not age. Therefore, while RDS is a promising method, some methodological improvements would need to be made for it to be useful for the ONS. With regards to the practical side of data collection, the tests showed that smartphones can be successfully used for this purpose. The test that targeted boat dwellers showed that members of the population of interest must interact sufficiently frequently as a network for RDS data collection to work.

Back to table of contents

2. The case for using respondent-driven sampling

Producing statistically robust profiles of small or “rarely heard” populations has always been a challenge to researchers. This has had negative consequences for those populations, as policy making and legislation is less efficient when there is little high-quality data.

In theory, the issue of small population size can be addressed by taking a sufficiently large sample of the general population that would yield enough responses from the population of interest to arrive at a statistically robust profile of it. But that is very costly and does not address the issue of “rarely heard” populations, where the target population is less able, or less willing, to respond to the survey than the rest of the population.

A possible alternative is recruiting survey respondents through non-probabilistic sampling, such as snowball sampling. This approach allows for obtaining a desired sample size, but at the same time entails losing the statistical benefit of having a probabilistic sample. For example, using probabilistic sampling means the sample is unlikely to be skewed by selection bias (as long as non-response is approximately random).

However, in snowball sampling, various biases are likely to occur, as explained in Respondent-Driven Sampling: A New Approach to the Study of Hidden Populations by Douglas D. Heckathorn. For example, participants may disproportionally recruit peers that are similar to them, and people who are well networked have a higher chance of being invited to the study than those who are socially isolated.

Respondent-driven sampling (RDS), developed by Douglas Heckathorn in the United States in the late 1990s, combines snowballing sampling with a dedicated statistical estimator that corrects for some of the biases that arise from the snowballing. Despite its apparent potential, it has not been widely used, possibly because the data collection, usually completed face to face, was lengthy and costly.

However, technological advances, particularly the invention of the smartphone, have allowed this method to be applied on a larger scale. In addition, there is a growing recognition of the disadvantages experienced by groups underrepresented in the data landscape. As discussed in our report on recommendations for being more inclusive in our data (PDF, 1,032KB), this underrepresentation results from a lack of high-quality data and declining response rates in traditional surveys, and is the reason that we are now exploring the potential of RDS.

While RDS is a method that was developed for surveying “rarely heard” populations, there is no theoretical reason why it should not be used to survey the general population.

Back to table of contents

3. Aims of the project

An Office for National Statistics (ONS) project team commissioned Dr Filip Sosenko from the University of Glasgow, who has previous experience of using smartphone-based respondent-driven sampling (SB-RDS), to carry out the tests. The tests were aimed at establishing whether SB-RDS could become a useful tool for the ONS via a series of real-life tests.

The fundamental research question for the trial was:

Is the population profile from data collected through SB-RDS close enough to the true profile to be useful for the ONS?

We were also expecting to learn about the practical solutions and arrangements of running SB-RDS in the different populations included in the trial. This would also answer one of our secondary research questions. We wanted to know the best practice for implementing the SB-RDS method to standardise future work that might utilise this method.

We decided to trial the method with three different populations, selected by the ONS with input from Dr Sosenko. These are:

  • Residents of the York City Council, aged 16 years and over
  • People identifying as Gypsy, Roma or Traveller, living in England and aged 16 years and over
  • People living on a boat for at least three months in a year, living in England and aged 16 years and over

Residents of York City Council, aged 16 years and over

The aim of this test was to find out how close the population profile from RDS is to the true profile from Census 2021.

York was chosen as it is a relatively well-defined urban location without sizeable settlements close by. This helps to reduce the likelihood of capturing respondents who live elsewhere but have commuted to York, which would contribute to over-coverage in the data.

To recruit seed participants in York, University of Glasgow sub-contracted a field researcher from a market research company. The recruitment took place in a public location in the centre of York.

People identifying as Gypsy, Roma or Traveller, living in England and aged 16 years and over

This test was included:

  • because this group is generally underrepresented in surveys using traditional methods
  • to better understand how the group might interact with a new approach

Seed recruitment for this test was undertaken with the engagement of two support organisations: Traveller Movement and Roma Support Group.

People living on a boat for at least three months in a year, living in England and aged 16 years and over

This group may be underrepresented in traditional surveys as it is less likely to appear in address-sampling frames.

Seed recruitment for this test was undertaken with the engagement of the boat dwellers support organisation, Waterways Chaplaincy.

Back to table of contents

4. Overview of the respondent-driven sampling methodology

Respondent-driven sampling (RDS) is a three-step research method, with unique features in each step, as described in Sampling and Estimation in Hidden Populations Using Respondent-Driven Sampling by Matthew Salganik and Douglas Heckathorn.

Firstly, in the design step, the researcher must identify the target population and whether the study design is likely to meet RDS assumptions. Of particular importance is the assumption that the population of interest is not divided into subgroups where respondents are outside each other’s social networks (for example, because of language, ethnicity or social class). If the target population is divided into such subgroups, this creates “bottlenecks” in social networks. Snowballing is likely to be affected by these subdivisions in the population, and the RDS statistical estimator would be unlikely to sufficiently correct for this. Also, the sample size should be calculated according to the confidence intervals for which the researcher is aiming.

In the second step – sampling and data collection – “seed” respondents are recruited by the researcher (this is the initial respondent that we will encourage to recruit additional respondents). Typically, their number varies between 5 and 15, as small numbers of seeds with long chains is the ideal. Each “seed” is assigned a unique identifier and asked to fill in the survey questionnaire.

Importantly, the survey must include a question about the number of people in the respondent’s personal network who fit into the target population. Usually, this is achieved by asking a question such as:

“In the past three days, how many friends, acquaintances and family members living in York did you talk to (in person, by phone, text, social media or email)?”

This information is later used by the RDS estimator to down-weight respondents who have larger than average social networks. This is because those with larger networks have a higher chance of being recruited than those with smaller networks.

Once the “seed” has filled in the survey questionnaire, they are rewarded with a £10 gift voucher. They are also given a few (typically two or three) unique invitation codes to distribute between their network in the target population. Those follow-up contacts who are recruited complete the survey questionnaire, receive their reward and further invitation codes. Those who recruited the follow-up contacts receive a secondary reward, usually smaller than the primary reward. In this trial, the primary reward was £10, and the secondary was £5. Data collection continues until the target sample size is reached.

Lastly, in the third step, data analysis, the researcher applies one of the dedicated RDS estimators to analyse the data. Estimators correct for some of the biases that occur in snowballing. The three primary sampling biases that occur in “snowball” sampling are as follows:

  • members of the population who have larger social networks are more likely to be invited to the survey than those with smaller social networks
  • some types of respondents may recruit more efficiently (at a faster rate, or more recruits on average) than other types of respondents
  • some types of respondents may disproportionally over- or under-recruit people similar to themselves (this is known as recruitment “homophily/ heterophily”)

RDS estimation

When survey participants are randomly sampled from a sampling frame, observations are independent. This is no longer the case when the sample is recruited using social networks, as in “snowballing”. Social networks are not random because people tend to cluster with other individuals who are, to some extent, similar in demographics and attitudes. This means that there is usually dependency between observations: the social characteristics of the recruiter are “informative” with respect to the social characteristics of the recruit.

This dependency can be neutralised by using the Markov chain estimator. Markov chain is a succession of “states” in which (a) there is dependency between the current state and the previous state, and (b) the current state is chosen at random from possible/ available states.

The fundamental idea underlying RDS is that snowballing-type recruitment for a survey, using social networks, is a Markov chain process: the “state” (characteristic) of the recruit is dependent on the “state” (characteristic) of the recruiter. For this to be a true Markov chain, however, the recruiter needs to choose at random from the pool of available “states”, that is, from the pool of their network contacts. If that is the case, we can estimate the probability (prevalence) of a given characteristic in the population using the Markov chain estimator. All that we need is to know who recruited whom to the survey and what their characteristics were: this will give us a picture of which “states” followed which. For example, if we are interested in the proportion of males in the target population, the Markov chain estimator will calculate the proportion of time the process visits the “male” characteristic in the long run and will calculate the probability of being male (while correcting for dependency).

When data are collected on two groups (for example, female and male at birth), the Markov chain estimator formula for the proportion of the population with a given characteristic is:

“Proportion of males among people recruited by females” is the same as “transitions Female-Male as proportion of all transitions starting from Females”.

Generalising this into any binary situation and coding “the feature is present” as 1 and “the feature is absent” as 0, we have:

And finally using the number of transitions N:

When there are more than two exclusive categories, the basic principles of RDS estimation remain the same, but the calculations are more complex. They involve solving an over-determined system using linear least squares.

An important element of RDS estimation is the size of the respondents’ social networks. If the two exclusive sub-groups under study – such as males and females – have different average network size, the transitions formula presented here would provide an incorrect estimate. The proportion of the group with a larger average network size would be overestimated. That is because people who have larger networks are more likely to be selected for the study (invited to participate) than people with smaller networks. To counter this bias, RDS adds the average network size to the formula, proportionally down-weighting the group with larger networks and upweighting the group with smaller networks, to the point where the difference in network size disappears:

This explains why every RDS survey contains a question about the size of the respondent’s social network.

The RDS estimator described here is called “RDS-I” or “Salganik-Heckathorn” (SH), after the names of its inventors. If RDS assumptions are met (see below for the main assumptions for RDS), it corrects for dependency between the recruiter and the recruit, as well as for the first two sampling biases: systematic differences in network size and systematic differences in recruitment efficiency.

Another popular RDS estimator is called RDS-II, or “Volz-Heckathorn” (VH). Unlike RDS-I, RDS-II does not correct for differences in recruitment efficiency, but it has lower variance. It therefore should be preferred over RDS-I when groups under study did not differ with respect to recruitment efficiency.

The last mainstream estimator, the “Successive Sampling” (SS) estimator (demonstrated in Improved Inference for Respondent-Driven Sampling Data article by Gile), is an extension of RDS-II. It is designed to provide more accurate estimates when the sampling fraction is large, but it requires knowing the population size.

All of the above estimators assume that recruitment is at random. One much less used estimator, demonstrated in Linked Ego Networks article by Lu, does not rely on this assumption but it requires asking survey respondents about the composition of their social network. This increases cognitive burden and lengthens the survey questionnaire. However, this estimator corrects for all three sampling biases mentioned earlier: systematic differences in network size, systematic differences in recruitment efficiency, and recruitment homophily/ heterophily.

Important RDS assumptions

There are a few assumptions that need to be met for this formula to work. An important one is that recruiters choose a recruit from their social network at random. (the Lu estimator does not make this assumption). If this assumption is violated, the result is a phenomenon called “recruitment homophily” or “recruitment heterophily”. Recruitment homophily occurs when recruiters are disproportionally likely to recruit people in their network who are similar to them. Recruitment heterophily occurs when recruiters are disproportionally likely to recruit people in their network who are dissimilar to them.

Another important assumption is that the population under study is not split into completely exclusive sub-populations. There needs to be at least some connection between sub-populations for RDS to work. If this assumption is not met, the recruitment would be “stuck” within the subpopulation(s) in which the recruitment started, and Markov estimation would provide incorrect results.

It is also assumed that after going through sufficiently many transitions the Markov chain will converge to its “equilibrium” regardless of the characteristics of the “seed” respondent, and as a result the estimate will stabilise around a certain point. Adding further observations may move the estimate to some extent, but it will keep going back to the estimate from the moment the equilibrium was reached.

Back to table of contents

5. Study methods

We used a smartphone-based respondent-driven sampling (SB-RDS) data collection system. The SB-RDS system was developed in 2019/2020 for a previous study, as outlined in the A methodological advance in surveying small or “hard-to-reach” populations article by Sosenko and Bradley.

Smartphone-based respondent-driven sampling process

The SB-RDS process works as follows. The participant receives a unique invitation code from a “parent” (someone recruited by us) and texts this code to a designated mobile number. Within a few seconds, the participant receives an automated text with a welcome message and a link to the survey. This link is specific to the code that was sent, so it cannot be reused or shared. The participant fills out the survey online on their smartphone. Within a few seconds of completing the survey, the respondent receives a text message with a reward (£10 electronic gift voucher) and another message with follow-up unique invitation codes for friends. When a “child” (someone recruited to the survey by the “parent”) completes the survey, the “child” receives a primary reward (£10 electronic gift voucher) and a secondary reward (£5 electronic gift voucher) is automatically texted to the “parent”.

With their permission, we geolocated respondents for the residents of York survey to ensure they were not outside the target population. These GPS data are not known to the researcher and are not saved in the database of survey responses.

This system includes three subsystems, which interact with each other.

System A

System A was responsible for sending automatic text messages to participants. It needs to be a programmable service, for example, one that can be programmed using bespoke computer code.

Our system used Twilio and the code was written in Javascript.

System B

System B was a database that consisted of two tables – one with valid invitation codes and another one with rewards (electronic gift vouchers, which take the form of a hyperlink).

In our system, the database was cloud-based (via Amazon Web Services) but it could be hosted on an institutional server.

System C

System C was responsible for administering the online survey. It required the following two functionalities:

  • to be able to retrieve respondent’s invitation code from the survey URL
  • to have the “end URL” option, which triggered a custom URL link at the end of the survey

The following three functionalities are technically not necessary, but are desirable:

  • triggering the phone’s GPS functionality and receiving the reading
  • setting a “prevent repeat participation” cookie
  • recording how long the respondent spent answering each question

Our system used LimeSurvey, which offers all these functions, but any online survey platform with the first two functionalities will work.

Preventing survey fraud

The architecture of the SB-RDS system has been designed to prevent survey fraud. Like other online data collection methods involving incentives, online RDS faces the multiple challenges: multiple participation, ineligible participation and poor quality of responses from eligible participants.

The SB-RDS system prevents multiple participation and ineligible participation through safeguards built into Systems A, B and C. System C (online survey) records both the total survey completion time and the completion time for each question, which allows for an informed evaluation of the extent to which poor quality responses was an issue.

Multiple participation

Multiple participation is when a participant tries to obtain more than one incentive by filling in the survey more than once. In the case of SB-RDS, a participant may try to fill in the survey more than once using the same invitation code. They may also try to use a follow-up invitation code that is supposed to be used by a contact, thereby impersonating a “child” respondent.

Ineligible participation

Owing to the nature of the online data collection method, respondent anonymity can result in three types of attempts at ineligible participation:

  • a respondent may try to respond by trying to use someone else’s invitation code or by trying to guess a code
  • a respondent may not meet the requirement of residing in the area being studied but respond anyway to receive vouchers
  • a respondent may not meet one or more of the demographic eligibility criteria (such as being aged over 16 years) but give false answers to give the impression that they do meet the age requirement

Poor quality of responses from eligible participants

The respondent may rush through the questionnaire, with little attention given to reading the questions and response options.

Back to table of contents

6. Data analysis

We applied the two most commonly used and most researched estimators: RDS-I and RDS-II. We also applied the successive sampling (SS) estimator in the residents of York survey, as we knew the population size from Census 2021 estimates.

In the Gypsy, Roma and Traveller survey data analysis, we also applied Lu’s estimator. This estimator is much less commonly used, but has the advantage of not relying on the random recruitment assumption.

Back to table of contents

7. Results

Test 1: Residents of York survey

Participants completed the survey between 24 September 2024 and 28 October 2024, when data collection stopped because of sufficient sample size. There were 827 valid responses to the survey, of which 807 were the “effective sample”. Before analysis, we discarded 20 responses that came from participants aged under 16 years because they were not part of the eligible population.

Of the six seeds recruited by the professional market research interviewer, one was fully unproductive (did not recruit anyone) and one recruited only one person, after which the chain died out. The remaining four seeds were productive, with the resulting recruitment chains containing between 66 and 337 people.

Comparison with Census 2021 benchmark

In this section, we investigate the extent to which the respondent-driven sampling (RDS) results are different from the Census 2021 benchmarks.

Age

Very few people aged 65 years and over participated in the RDS survey, so we only show results for those aged 16 to 64 years to maximise comparability. We assume the lack of participants aged 65 years and over is, in part, because of lower smartphone usage in these age groups. This must be considered in future work to be more inclusive of this age group.

RDS estimates are quite different to the Census values for each age band, with none of the 95% confidence intevals (CI) containing the Census value (Table 2).

Sex

The distribution of sex also differed, though the 95% CIs managed to capture the true value (Table 3).

The three most commonly used RDS estimators, including the successive sampling (SS) estimator, all assume random recruitment. The lack of accuracy in the York estimates motivated us to investigate the extent to which recruitment may have been non-random. While formal diagnostics of non-random recruitment cannot be conducted using standard RDS data alone, it is still possible to make some informative inferences from the data we collected.

We asked a subset of respondents about the sex composition of their personal networks. This data showed that the proportion of men within men’s networks was similar to the proportion of women within women’s networks. So if the recruitment had been at random, we would expect the proportion of men recruited by men to be similar to the proportion of women recruited by women. However, the observed proportions differed substantially. While most people recruited by women were women, most people who were recruited by males were also female.

Recruitment also appeared non‑random with respect to ethnicity. Transitions from non‑Black to Black participants occurred 3.5% of the time. This is compared with an expected 1%, if recruitment was random, given York’s 1% Black population.

Test 2: Gypsy, Roma and Traveller survey

Participants completed the survey between 17 January 2025 and 11 February 2025, when the recruitment chains died out (though the survey window was open until the end of March). There were 346 valid responses to the survey, of which 222 were the effective sample from the target population. Before analysis, we discarded 80 responses from participants who were:

  • aged under 16 years
  • aged 75 years and over
  • not living in England

There were 44 responses that appeared to come from a fraudster using multiple SIM cards or virtual phone numbers, because of the series of parent-child responses from very similar phone numbers.

We provided support agencies with 16 invitation tokens for seeds, of which six were used. However, only two seeds were substantially productive, while four recruited between one and three people, after which their chains died out. Of the two productive chains, one had 115 people and the other had 97 people. The number of recruitment waves was 35 in the former and 11 in the latter.

Comparison with Census 2021 benchmarks

We are aware of evidence for the underparticipation of people identifying as Gypsy, Roma and Traveller. We present these RDS results, together with Census 2021 results, for comparison. Our RDS survey collected data mainly from people identifying as English Gypsy/Travellers or Irish Travellers, rather than Roma. The Census data that we use is also from these two groups (in the ethnicity category “Gypsy or Irish Traveller”). The census data used is for England only and not Wales.

Sex

Sex is the only characteristic where the proper benchmark is available, suggesting results should be split around 50/50. RDS estimators relying on at-random assumption severely underestimated the proportion of women, while the Lu estimator (which relies on network composition information) produced only a minor underestimate (Table 4).

Age

Table 5 shows the remaining characteristics, alongside Census 2021 data, for comparison.

Test 3: Survey of boat dwellers

Participants completed the survey between 3 March 2025 and 28 March 2025, when the survey was deactivated. There were 35 valid responses to the survey, coming from 12 seeds. Five seeds were unproductive, in that they completed the survey but did not manage to recruit anyone. The most productive chain had 10 participants and five recruitment waves.

At the end of the data collection, we discussed the process with our support contact. We learned that boat dwellers interact with each other much more when the weather is warm, so it would be better to attempt surveying in the summer months to help the spread of invitations. We also learned that some boat dwellers have poor, or no, mobile data reception.

Scale of attempted survey fraud

The scale of attempted survey fraud has been similar to that observed in a previous SB-RDS study. A substantive number of participants attempted to participate more than once to claim an additional reward. A smaller but still non-negligible number of people tried to participate without being invited to the survey. All but one type of fraud was prevented by the SB-RDS system.

Someone tried to exploit a gap in the system by using virtual mobile numbers or multiple SIM cards in the Gypsy, Roma and Traveller survey. In a positive development, these were spotted by the almost sequential set of phone numbers that appeared in a chain, which alerted us to a vulnerability in the system that we subsequently patched. It is positive that this happened in a test situation and not during a real survey. There are no remaining vulnerabilities that would allow large-scale survey fraud, as far as we can tell. We checked York responses retrospectively for this and identified no such attempts.

Unlike the geolocation used in the York survey, the Gypsy, Roma and Traveller and boat dweller surveys did not employ a mechanism for verifying eligibility. This means we cannot comment on the scale of attempted survey fraud from people who obtained a valid invitation token but who did not belong to the target population in those two surveys. We decided that it would not be appropriate to use screening questions to validate membership of either target population on this occasion.

We have found that a non-negligible proportion of respondents to the Gypsy, Roma and Traveller survey completed the questionnaire in a very short time. We suggest that this is addressed in the future, both retrospectively, by excluding such responses or carrying out a sensitivity test without them, and proactively.

Back to table of contents

8. Future developments

Estimation

Respondent-driven sampling (RDS) estimators that assume random recruitment are too inaccurate for real-world use and should be avoided, according to the results of the York test and our investigation into the extent to which recruitment may have been non-random.

In contrast, the Lu estimator was the closest to the true population proportion in the Gypsy, Roma and Traveller test, which suggests that future efforts should consider using it, at least as a backup option. However, more testing of its accuracy in real-life applications is needed.

Eligibility validation

Consideration was given to using screening questions to identify and exclude people who are not part of the target population for the Gypsy, Roma and Traveller and boat dweller surveys. However, we would not recommend this approach.

There is no straightforward solution when it comes to eligibility screening. Approaches that may work for one survey are not necessarily suitable for another.

One option would be to use "hard" screening questions that only members of the target population could answer. However, this risks creating the impression that respondents are not trusted to respond honestly. In addition, it is difficult to be confident that any set of screening questions would reliably distinguish eligible respondents from ineligible ones, without also inadvertently excluding genuine members of the target population.

An alternative would be to use "soft" screening approaches, such as asking respondents to provide a short written or audio response to a qualitative question. While this may deter some ineligible respondents, it would introduce additional burden for participants and require subjective assessment of responses. This could raise concerns around consistency, transparency and fairness in determining eligibility.

Given these limitations, we do not recommend relying on screening questions as the primary mechanism for determining eligibility. If concerns remain about ineligible participation, alternative methods for improving sample quality should be considered, particularly where options such as geolocation are not appropriate.

Experimenting with the level of incentive

It would be helpful to have more methodologically focused evidence around the level of incentive. It could be tested whether £5.00 or £7.50 would be enough as the primary reward if the survey questionnaire is short, and whether £10.00 is enough if the survey is long. Saving money is an obvious advantage of setting incentive levels low. However, if the incentive is low, it may cause recruitment chains to break and it may result in severe under-participation of citizens with higher incomes. Estimation may not be able to address such bias effectively. Higher incentives increase costs but also increase the possibility of people committing survey fraud to receive the incentive.

Future use of smartphone-based respondent-driven sampling

RDS is potentially a promising method for surveying populations without the use of probabilistic sampling. It is particularly valuable where the study population is unwilling or unable to take part in traditional household surveys or censuses.

This project demonstrated both the potential and current limitations of smartphone‑based respondent-driven sampling (SB‑RDS) as a data collection approach for rarely heard populations. Across the three pilots, SB‑RDS proved operationally feasible, as recruitment chains were formed successfully in two tests and fraud‑prevention mechanisms were largely effective. The method therefore shows clear promise as a scalable, lower‑cost alternative to traditional field‑based RDS or address‑based sampling for populations that are difficult to identify through standard frames.

However, the accuracy of population estimates produced through SB‑RDS varied across tests.

For the York general‑population pilot, RDS estimates aligned reasonably with Census 2021 for sex. However, age distributions diverged substantially, with no 95% confidence intervals for age bands overlapping Census values. This suggests that SB‑RDS may struggle to generate representative estimates for characteristics strongly associated with smartphone use and digital literacy without further methodological development.

For the Gypsy, Roma and Traveller test, the study demonstrated that reliance on the standard “at‑random recruitment” assumption produces markedly biased estimates. In contrast, the estimator that exploits network composition information performed substantially better. These findings reinforce the need for us to move beyond conventional RDS‑I and RDS‑II estimators, and to further explore approaches that explicitly adjust for non‑random recruitment behaviour or that do not rely on the random recruitment assumption altogether. However, this involves increased burden because of additional network composition questions are required.

The boat‑dwellers test further highlighted the importance of contextual factors for successful recruitment. Seasonal patterns in mobility and interaction, and variable access to mobile data, had a direct impact on chain depth and sample size. These insights emphasise that SB‑RDS requires careful tailoring to population characteristics, timing and engagement strategy.

Across all three tests, survey fraud emerged as a noteworthy operational challenge, which is consistent with previous online studies. Most fraudulent attempts were successfully blocked. However, the experience underlines the need for:

  • robust eligibility checks
  • proactive data quality safeguards
  • a review of incentive structures

Without appropriate screening, the validity of estimates for small populations cannot be assured.

Overall, the project provides evidence that SB‑RDS has the potential to become a valuable addition to our methodological toolkit, particularly for populations that cannot be reached effectively through address‑based sampling.

Back to table of contents

9. Cite this page

Office for National Statistics (ONS), released 28 August 2026, ONS website, supporting methodology article, Testing smartphone-based respondent-driven sampling

Back to table of contents