MI MOHAMUD IBRAHIM · FIELD FILE ← Exit to Articles
File No. 027 — Horn of Africa MEAL Series

The Synthetic Respondent

A Provenance-First Framework for Detecting AI-Assisted Fabrication in Humanitarian Field Data

Why the sector's instinctive defence — the AI-text detector — should not be treated as evidence of fabrication, and what a defensible, integrity-focused alternative looks like in low-access settings.

Abstract

The integration of generative artificial intelligence into field data workflows does not create a new category of research fraud so much as lower the cost, and raise the plausibility, of an old one: the fabrication of interviews that never took place, long known in survey methodology as curbstoning.1 This paper argues that the sector's most intuitive response — submitting suspect text to an automated AI-content detector — is not a sound basis for any integrity finding in this operating context. Published evidence indicates such detectors misclassify non-native English writing as machine-generated at high rates,5 which describes the prose of precisely the enumerators, translators, and field officers a MEAL system depends upon. The argument here is about the evidentiary status of these tools, not a claim that no future detection technology could have any use: AI-text detectors should not be treated as reliable evidence of fabrication, and in particular should not form the basis of disciplinary, exclusion, or misconduct decisions.

I propose a reframing. The operative question is not "does this text look AI-generated?" but "can this record be credibly traced to a real respondent, at a real time and place, through a defensible collection process?" The paper separates four questions that are routinely conflated — AI involvement, interview authenticity, data fidelity, and fabrication — and sets out a five-layer, provenance-first protocol, an inference ladder distinguishing suspicion from established fabrication, an operational decision framework, and the ethical safeguards without which such a protocol becomes a tool of injustice against national staff.

Keywords — data fabrication · curbstoning · generative AI · research integrity · data provenance · paradata · back-checking · localisation · Somalia · Horn of Africa
§ 01 / 07

An old fraud, newly cheap to commit

Long before large language models, methodologists had a name for the interview that never happened. Curbstoning — an enumerator sitting on the curb and inventing plausible responses rather than knocking on the door — has been documented across professional survey operations for decades, and the empirical literature indicates it is neither vanishingly rare nor confined to unprofessional teams.1 In one large South African survey, forensic re-analysis concluded that a meaningful share of interviews had been fabricated by fieldworkers despite standard controls — a finding from that study and context, not a universal rate, but a caution against assuming one's own data is immune.4 Separately, an audit of many widely used public-opinion datasets reported that a substantial minority contained exact or near-duplicate records above a five-percent threshold.2 Fabrication, on this evidence, is a recurring risk in field research under pressure rather than an exotic failure.

Generative AI does not introduce this problem; it changes its economics. What previously required a fabricator to imagine internally consistent answers — cognitively taxing, and historically detectable partly because human fabricators tend to under-produce variance3 — can now be produced quickly, in volume, with fluent narrative texture in the open-ended fields that once resisted fabrication most. The historical tell of the fabricated qualitative response — flat, repetitive, thin — is precisely the weakness a capable language model can mask.

Four questions that are not the same question

Most confusion in this area comes from collapsing four distinct questions into one. Keeping them separate is the analytical spine of everything that follows, because a record can trigger one without triggering the others.

A · AI Involvement

Was a generative model used?

At any point in producing the record. On its own, this is the least consequential question — and the one detectors purport to answer.

B · Interview Authenticity

Did the encounter occur?

Was there a real respondent, in a real place, at the stated time? This is a question about events, not text.

C · Data Fidelity

Is the record faithful?

Does the stored entry accurately represent what the respondent actually said or did — without silent paraphrase or alteration?

D · Fabrication / Falsification

Was it invented or altered?

Was information fabricated, materially changed, or falsely attributed to a respondent? This is the integrity violation that matters.

These are orthogonal. A record may have AI involvement (a translator used a model, and disclosed it) while remaining authentic, faithful, and non-fabricated. Conversely, a record may have zero AI involvement and still be pure fabrication — classic curbstoning. It follows that detecting AI involvement is neither necessary nor sufficient for establishing fabrication. The protocol proposed here therefore targets B, C, and D — authenticity, fidelity, and integrity — and treats A as, at most, a weak and easily-mistaken signal.

Why this matters acutely in Somalia and similar settings — and why now

The risk concentrates wherever three conditions coincide: remote management, high enumerator turnover, and thin independent verification. In the author's evaluation practice across Banadir, Puntland, Jubbaland, Galmudug, South-West State and Somaliland, all three are structural rather than incidental: international staff frequently manage programmes remotely, so a large share of collection runs through local enumerators and partners; access is uneven and security-contingent, so a proportion of sampled locations cannot be independently re-visited; and turnover means the person who collected a figure is often gone before anyone asks how. Somali-to-English translation adds a further layer, since much narrative data is rendered into English by field staff for whom it is a second or third language — the very population, as Section 02 shows, that AI-text detectors are most prone to misjudge.

These conditions are tightening, though the evidence varies in strength and should be read at the right level. At the sector level — across 22 monitored humanitarian operations, not Somalia specifically — the OCHA Centre for Humanitarian Data estimates that the share of crisis data assessed as available and up-to-date in its HDX Data Grids fell from 74 percent to 68 percent between its 2025 and 2026 assessments, a decline it attributes largely to the contraction in humanitarian financing.9 Beyond that primary figure, secondary sector analysis, aggregating agency reports, describes reductions in the in-person verification capacity that underpins data quality, including disbanded enumerator networks; some Somalia-specific figures — such as a sharp contraction in the number of locations where routine collection remains feasible — appear in that analysis only as reported projections, which I have not been able to trace to a primary agency publication and therefore treat as indicative rather than established.11 The overall direction, however, is not seriously in doubt: the verification layer is thinning at the same moment the tooling for fluent fabrication has become increasingly accessible and capable. That coincidence — rising opportunity, falling scrutiny — is why this is a present problem, not a future one.

Working definition AI-assisted fabrication, in this paper, is any record entered into a monitoring or research dataset as if collected from a real respondent, where a generative model has been used to invent, materially alter, or falsely attribute content — spanning fully synthetic interviews and AI-authored narrative grafted onto otherwise real records. It is deliberately distinguished from disclosed AI assistance (e.g. logged translation or cleaning that preserves the original), which is a governance question, not misconduct.
§ 02 / 07

Why the obvious defence is not a defence

Confronted with suspected AI text, the reflexive move is to run it through an AI-content detector and treat the score as a verdict. This paper's central methodological claim is narrow but firm: whatever the future of detection technology, today's stylometric detectors do not have the evidentiary reliability to support an integrity finding in this context — and using them to accuse, discipline, or exclude field staff is both unsound and unjust.

False-Signal-01

Detectors tend to penalise the writing of your own field staff

These tools infer machine authorship from statistical properties of text — chiefly perplexity (how predictable each word is) and burstiness (how much sentence length and structure vary), with more recent curvature-based refinements.6 Machine text tends to score low on these measures. So, however, does the careful, grammatically regular, lower-variance English of a competent non-native writer.

What the evidence shows In a controlled study by Liang and colleagues, evaluating seven detectors on 91 TOEFL essays by non-native English speakers and 88 US eighth-grade essays, a majority of the non-native essays were misclassified as AI-generated (a mean false-positive rate of 61 percent), and 97.8 percent were flagged by at least one detector, while the same tools were near-perfect on the US essays.5 The sample is small and drawn from an education context, so the figures should be read as indicative of a direction, not as a precise rate; the relevant read-across is the population — non-native English writers — not the setting. A separate independent evaluation of fourteen detection tools concluded they are "neither accurate nor reliable," biased toward labelling machine text as human, and — directly relevant here — found they cannot reliably cope with text translated from another language.7 The developer of a widely used chatbot withdrew its own classifier, reporting it correctly identified only around a quarter of AI text while falsely flagging roughly one in eleven human passages.8
In a Somali MEAL system, the writers whose English is most regular — enumerators trained to a house style, translators rendering Somali into clean English — are among the writers such a detector is most likely to flag. A positive result may reflect non-native fluency rather than fabrication, and by itself is not a reason to suspect anyone.

The consequence is both epistemic and ethical. Epistemically, a tool with these error characteristics cannot by itself support a defensible finding: a positive result is ambiguous, and a negative result is easily manufactured, since light manual editing or paraphrasing of machine text has been shown to push a large share of it past these tools undetected.7 Ethically, building quality assurance around such a tool means constructing a process whose false accusations fall hardest on national staff — reproducing, in miniature, the asymmetry of trust that the localisation agenda exists to dismantle.

What this argument is, and is not It is not a claim that fabrication cannot be detected, nor that detection tools can never improve. It is a claim that fabrication cannot be established by asking whether the text looks synthetic. The dependable signal is not in the prose. It is in the record's provenance — where it came from, when, through what process — and in whether a second contact with the respondent corroborates it.
§ 03 / 07

Four vectors by which AI text enters a dataset

Not all contamination is equal, and conflating the cases produces both false alarms and missed fabrication. Two of these vectors are integrity violations; two can be honest practice that nonetheless corrupts the evidentiary record when undocumented. Each maps onto the four questions of Section 01.

Vector-01

The wholly synthetic interview

A record generated end-to-end by a model and submitted as a completed interview. There is no respondent, no visit, no consent. It typically appears in batches, produced to close the gap between a target sample and what a team actually reached.

Classification — fabrication (fails B, C, and D). Straightforward misconduct.
Vector-02

AI-authored narrative on a genuine record

A real respondent was interviewed and the structured fields are real, but the narrative fields — quotes, the "why," the case story — are invented by a model rather than derived from what the respondent actually said. The distinction that matters is meaning, not tool use: this vector is falsification only when the model supplies content the respondent did not express, or attributes to them words they did not say. It is not this vector when a model is used to transcribe, translate, summarise, tidy grammar, or apply structured coding to the respondent's genuine account, the transformation is disclosed, and the original is preserved — that is Vector 04. The test is whether the respondent's meaning survives intact and traceable.

Classification — falsification / serious integrity violation (passes B, fails C and D). The respondent's account is displaced by an invented one.
Vector-03

Machine translation presented as verbatim

A respondent spoke Somali; the response was rendered into English by a model and stored as if it were a faithful, word-for-word translation. Often well-intentioned — but it silently substitutes the model's paraphrase for the respondent's words, flattening idiom, hedging, and the ambiguity a qualitative analyst needs.

Classification — not inherently misconduct; a fidelity/provenance problem (risk to C) when undocumented or misrepresented as verbatim. Defensible when disclosed and the source-language response is preserved.
Vector-04

AI-assisted cleaning, summarising, or polishing

Real responses run through a model to transcribe, fix grammar, standardise, or summarise before analysis. This is already happening in the sector: KoboToolbox, for instance, deployed AI-assisted voice capture, transcription and translation during the early-2026 Mozambique floods response, built on a "human-in-the-loop" design that requires staff to review and verify each AI-generated transcript and translation before the data is used.12 That review-and-preserve design is precisely what keeps such assistance on the right side of the line — so a blanket prohibition on any AI contact is neither realistic nor, in itself, warranted.

Classification — can be legitimate practice when disclosed, controlled, and the original response is preserved (protects C). Corrupting only when hidden.

The taxonomy yields the governing principle of this paper: the issue is not whether AI touched the data, but whether the transformation is disclosed, traceable, faithful, and consistent with the research protocol. Vectors 01 and 02 are integrity violations to be caught and addressed through formal procedures; Vectors 03 and 04 are process risks to be governed by disclosure and by preserving the source. A system that cannot tell these apart will either wave through fabrication or wrongly treat honest, efficient practice as misconduct.

§ 04 / 07

A provenance-first detection protocol

If the dependable signal is not in the text, it is in the record's provenance and in the respondent's corroboration. The five layers below move from cheapest and most automatable to most resource-intensive, and are designed as routine data-quality assurance rather than as tools reached for only after suspicion. None is individually conclusive; their value is cumulative and triangulated. The pre-AI fabrication literature approached detection through multiple, convergent statistical indicators paired with targeted re-interviews rather than any single test;1 the stronger position that no single signal should ever be treated as decisive on its own is this paper's methodological conclusion, and it governs everything that follows.

L1

Metadata & paradata forensics

Interrogate what the collection tool records automatically. The available signals differ by platform and configuration, so this must be stated conditionally rather than assumed.

Commonly available where configured: GPS coordinates, submission timestamps, and interview duration. Available on some platforms or only when enabled: device identifiers, and process paradata such as field-level timing, edit history, or paste events.10 Sometimes unavailable entirely. The protocol should be scoped to what a given deployment actually captures.

Signals — where the platform exposes such process metadata: implausibly short or uniform durations; GPS points clustering at one location or absent; batches submitted from a single device outside field hours; narrative fields pasted rather than typed.
Limitation: these are indicators of field presence, not proof. GPS, timestamps, and device data can be manipulated or spoofed, and must be read alongside other evidence, never alone.
L2

Duplication & near-duplication analysis

Compute pairwise similarity across records to surface interviews too alike to be independent — the established "percent-match" family of tests for cut-and-paste and templated fabrication.2

Signals — batches that, despite surface fluency, share latent structure; recycled records; partner submissions that duplicate one another.
Limitation: a model can be prompted for high diversity, raising the floor of detectable similarity; and some genuine similarity is expected in short, closed-response instruments. Necessary, but not sufficient alone.
L3

Distributional signatures — where theoretically appropriate

For numeric fields, compare observed distributions against those expected on theoretical grounds: reduced response variance relative to genuine interviews,3 planted questions whose true distribution the analyst knows but a fabricator does not, and — only where the data-generating process makes it appropriate — digit-frequency tests such as Benford's Law.

Signals — numeric patterns that depart from a well-justified expectation, prompting closer inspection of the associated records.
Limitation: Benford-type tests apply only to certain naturally-occurring, multi-scale quantities; many legitimate field variables do not follow a Benford distribution, and non-conformance is not, by itself, evidence of fabrication. Applies to structured fields, not narrative.
L4

Independent back-checks & audio audits

Re-contact a random sub-sample of respondents to confirm the interview occurred and that key answers match; where consented audio was captured, audit-listen a sample against the recorded transcript.4 This is the layer that most directly addresses authenticity (question B).

Signals — a graded outcome, not a binary: failed contact (respondent unreachable), contradictory back-check (reached, but answers or the encounter do not match), confirmed non-occurrence (respondent states no interview took place), and corroborated fabrication (non-occurrence plus supporting evidence).

Limitation: a failed back-check must not be read as fabrication. Respondents are legitimately unreachable for many reasons — changed numbers, movement, insecurity, network gaps, incorrect contact details, or simple refusal. Failed contact raises concern and warrants follow-up; only confirmed non-occurrence with corroboration approaches proof. This is also the layer most degraded by the current access and funding contraction — and the most decisive, which is the central paradox.
L5

Process governance & disclosure-by-design

Shift part of the burden upstream: require preservation of the original-language response and any raw audio, mandate declaration of any AI assistance in translation or cleaning, and log tool use as part of the record — turning Vectors 03–04 from hidden corruption into documented, auditable transformation.

Effect — makes honest-but-corrupting cases visible rather than punishable, and narrows the space in which undocumented text is acceptable, thereby also deterring Vectors 01–02.
Limitation: governs behaviour; does not, by itself, detect a determined violator. It is the frame within which L1–L4 operate.

From indicators to claims: an inference ladder

The protocol's honest output is a position on a ladder of claims, not a verdict. Conflating these rungs is the same error the AI detectors commit — presenting a probabilistic signal as a categorical fact. Each step up requires more, and more independent, evidence.

1AI may have been involved. The weakest and least consequential claim — and, on current tools, the least reliable. Rarely actionable on its own.
2The record has suspicious provenance. One or more metadata/duplication/distributional anomalies. A reason to look closer, not to conclude.
3The interview may not have occurred as claimed. Provenance anomalies plus a failed or contradictory back-check.
4The respondent's account appears materially altered. Evidence that stored narrative diverges from a preserved source (audio, original-language text).
5Fabrication or falsification has been established. Confirmed non-occurrence or material alteration, corroborated by independent evidence. Only at this rung does the evidence support a misconduct conclusion — but whether one follows also depends on organisational policy, the applicable evidentiary standard, due process, and the accused person's right of response, none of which this protocol can supply.

Note that the protocol never needs to name the technology an enumerator did or did not use. It assesses record integrity and provenance — questions B, C, and D — and is deliberately agnostic about question A, because the invisible presence of a model is neither knowable with confidence nor, in itself, the thing that matters.

§ 05 / 07

What a MEAL manager should actually do

A framework that stops before the decision is of little use in the field. The following tiers translate the inference ladder into proportionate action. They deliberately avoid arbitrary numeric cut-offs; the organising principle is that multiple independent indicators raise confidence, while no single anomaly establishes fabrication. "Independent" is doing real work here — three anomalies that all stem from the same cause (a misconfigured device clock, say) are one indicator, not three.

Low concern · a single weak anomaly, no corroboration

Document the observation and retain the record. Note it for pattern-tracking across the dataset. No action against any individual.

Moderate concern · multiple anomalies, or one strong one

Flag the affected records and conduct targeted verification — prioritised back-checks, a closer look at the associated enumerator's wider submissions, and review of preserved source material.

High concern · multiple independent indicators converge

Quarantine the affected records pending investigation so they do not flow into reporting, and escalate through the agreed data-integrity process. Records are held, not yet judged.

Confirmed integrity breach · ladder rung 4–5, corroborated

Follow the organisation's formal investigation and, where warranted, exclusion or misconduct procedures — with due process, documentation, and a right of response for anyone accused.

Illustrative scenarioHypothetical — invented to show how the layers interact. The numbers below are illustrative, not findings from any study or real dataset.

A food-security endline returns 400 household records from four districts. Routine QA surfaces, in one enumerator's batch of 60: interview durations averaging under four minutes against a survey norm nearer twenty (L1); a cluster of GPS points falling on a single roadside coordinate rather than dispersed across households (L1); and open-ended "coping strategy" responses with unusually low variance and near-duplicate phrasing across supposedly unrelated respondents (L2–L3).

No one of these proves anything. Short durations can mean an efficient enumerator or a short-form module; GPS clustering can mean a densely settled IDP site; low narrative variance can mean a genuinely uniform experience. But three independent indicators converging on the same 60 records moves the assessment to "high concern" — quarantine and targeted back-checks (L4), not accusation.

The back-checks then decide the matter. If respondents are simply unreachable, concern remains but nothing is established. If several confirm no interview took place, the assessment climbs to corroborated fabrication and enters formal procedures — while the other 340 records, and this enumerator's other work, are judged on their own evidence, not by association.

§ 06 / 07

An anomaly is a reason to investigate, not a reason to accuse

A provenance protocol concentrates scrutiny on the people with the least power in the system: national enumerators, translators, and field officers, often employed on short contracts and managed remotely by international staff. Deployed carelessly, it becomes an instrument of exactly the injustice this paper set out to avoid. The following safeguards are not an appendix to the method; they are part of it.

Guard against the non-native-English trap at every layer, not just the detector. The same bias that makes AI-text detectors unreliable can seep into human judgement — treating fluent-but-regular English, or "too clean" narrative, as inherently suspect. Fluency is not evidence. Anomaly signals should be tied to provenance and occurrence, never to how the prose reads.

Presumption of good faith, and due process. A flag is a question, not a charge. Anyone whose work is investigated should know the concern, see the basis for it, and have a genuine opportunity to explain before any adverse decision — with the recognition that power asymmetry makes it easy for a national staff member to be presumed guilty and hard for them to contest it.

Proportionality. The intensity of investigation should match the strength and independence of the indicators, not the convenience of blaming the nearest contractor. Quarantining records is reversible; a misconduct finding is not.

Privacy and consent in the forensic layers themselves. GPS traces, device identifiers, and audio recordings are personal data. Collecting and scrutinising them for integrity purposes must rest on informed consent, clear retention limits, secure handling, and applicable data-protection standards — the QA process must not itself become a surveillance or protection risk for respondents or staff. Audio audits in particular require that respondents consented to recording for this use.

The asymmetry to keep in view A false negative lets one fabricated record through. A false accusation can end a local professional's livelihood and reputation in a labour market with few alternatives. These costs are not symmetric, and a defensible protocol is calibrated accordingly: slow to accuse, quick to verify, and willing to hold uncertainty as uncertainty.
§ 07 / 07

The stylometric reflex versus the provenance discipline

Stylometric Reflex
  • Asks: does this text look AI-generated?
  • Evidence lives in the words themselves
  • Error concentrates on non-native writers — your own staff
  • A negative is easily manufactured with one edit
  • Presents a probabilistic score as a verdict
  • Treats any AI contact as if it were misconduct
Provenance Discipline
  • Asks: can this record be traced to a real person, place, and moment?
  • Evidence lives in metadata, corroboration, and preserved sources
  • Scrutiny attaches to records, with safeguards for the people behind them
  • A fabricated respondent cannot be corroborated on re-contact
  • Places each claim on an explicit ladder from suspicion to proof
  • Separates fabrication from disclosed, faithful assistance

The argument reduces to a single reorientation. The arrival of generative AI in field research is widely received as a novel detection challenge demanding a novel detection tool. It is better understood as a stress test of a discipline the sector already possessed and had let atrophy: maintaining an unbroken chain of provenance from a reported figure back to a real person in a real place. The fabrication literature that predates the language model already showed how this is done — through metadata, duplication tests, appropriate distributional checks, and re-contact.1 The conclusion this paper draws from it is that such work must be done by triangulation, holding any single signal as a question rather than an answer.

What generative AI changes is the cost of the two failures. It makes fabrication cheaper to commit; and, through the biased detector, it makes false accusation cheaper to level against the very national staff the system depends on. A serious response refuses both. It declines to treat fluent text as evidence of a real interview, and equally declines to treat fluent text as evidence of a fake one. It rebuilds, instead, the unglamorous infrastructure of provenance — and pairs it with the safeguards that keep an integrity process from becoming an injustice.

Closing note An evaluation's credibility was never a property of how its testimonies read. It was a property of whether a real person, in a real place, actually said them — and whether the system can still show its work. The synthetic respondent is a threat only to a system that had stopped asking that question. The remedy is not a better detector. It is the discipline of being able to answer it — carefully, proportionately, and with the benefit of the doubt extended to the people who do the fieldwork.

Sources & evidentiary scope

Figures are reported as stated in the cited works, with their original populations and contexts preserved. Scope tags indicate whether a source speaks to a general/methodological question, a regional or sector-wide trend, or a Somalia-specific claim. Somalia-specific operational conditions not carrying a citation are drawn from the author's evaluation practice and are framed as practitioner observation, not measured fact.

  1. Koczela, S., Furlong, C., McCarthy, J., & Mushtaq, A. (2015). Curbstoning and beyond: Confronting data fabrication in survey research. Statistical Journal of the IAOS, 31(3). Method
  2. Kuriakose, N., & Robbins, M. (2016). Don't get duped: Fraud through duplication in public opinion surveys. Statistical Journal of the IAOS, 32(3). Near-duplicate ("percent-match") detection; a substantial minority of surveys examined contained excess near-duplicate records. Method
  3. Bredl, S., Winker, P., & Kötschau, K. (2012). A statistical approach to detect interviewer falsification of survey data. Survey Methodology. Low-variance and Benford-type signatures, developed on German Socio-Economic Panel data. Method
  4. Finn, A., & Ranchhod, V. (2017). Genuine fakes: The prevalence and implications of data fabrication in a large South African survey. The World Bank Economic Review, 31(1). Prevalence figure is specific to that survey and context. South Africa
  5. Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., & Zou, J. (2023). GPT detectors are biased against non-native English writers. Patterns (Cell Press), 4(7). >60% of non-native TOEFL essays misclassified; 97% flagged by at least one of seven detectors; near-perfect accuracy on native-speaker essays; misclassification reduced by a fluency-enhancing edit. Population read-across (non-native writers), not humanitarian setting. Method
  6. Mitchell, E., Lee, Y., Khazatsky, A., Manning, C. D., & Finn, C. (2023). DetectGPT: Zero-shot machine-generated text detection using probability curvature. Proceedings of ICML 2023. Cited as representative of the perplexity/curvature detector family. Method
  7. Weber-Wulff, D., Anohina-Naumeca, A., Bjelobaba, S., et al. (2023). Testing of detection tools for AI-generated text. International Journal for Educational Integrity, 19(26). Evaluated 14 tools and concluded they are "neither accurate nor reliable," with a bias toward classifying AI text as human and degraded performance under paraphrasing, obfuscation, and translation. Method
  8. OpenAI (2023). Public statement on the discontinuation of its AI-text classifier, citing low accuracy (reported ~26% true-positive on AI text; ~9% false-positive on human text). Method
  9. OCHA Centre for Humanitarian Data (2026). The State of Open Humanitarian Data 2026. 68% of crisis data available and up-to-date across 22 humanitarian operations, down from 74% the prior year. Sector-wide across 22 operations — not a Somalia-specific figure. Sector-wide
  10. Kreuter, F. (Ed.) (2013). Improving Surveys with Paradata: Analytic Uses of Process Information. Wiley. Foundational treatment of paradata; availability of specific paradata varies by platform and configuration. Method
  11. Secondary sector analysis (e.g. PRISM, The State of Humanitarian Data, 2026), drawing on OCHA, IOM, UNHCR and WFP figures, reporting reductions in verification capacity and disbanded enumerator networks, with some Somalia-specific coverage projections. Reported/projected figures via a secondary aggregator; treat as indicative, and regional except where explicitly Somalia-specific. Regional / Somalia proj.
  12. KoboToolbox / Cisco (2026). Reporting on AI-assisted voice capture, transcription and translation with a human-in-the-loop review model, piloted by KoboToolbox during the early-2026 Mozambique floods response (a non-Somalia deployment; the flood event is independently corroborated by contemporaneous agency reporting). Source is the platform and its corporate partner; cited as an illustration that AI is entering the collection pipeline, not as a load-bearing empirical claim. Regional / vendor-reported