Skip to content

Data and constants ​

Loading the situation-report data, freezing it to a past cut-off, and the fixed quantities the model treats as known. load_observations reads data/observations.toml and is the entry point every fit starts from. The constants carry the province geography, the travel volumes and the sampler settings.

Index ​

Reference ​

BVDOutbreakSize.BASELINE_FIT Constant
julia
BASELINE_FIT

Value of a fit column identifying a row as the persistence baseline, which is no model's output.

source
BVDOutbreakSize.CHAMLA_CONFIRMED_CENTRAL Constant
julia
CHAMLA_CONFIRMED_CENTRAL

Central-scenario cumulative laboratory-confirmed case projection of Chamla et al. (2026) (WHO Regional Office for Africa, Lancet Infectious Diseases), as (date, median, lower_90, upper_90) tuples from their Table 1 (mean accepted R₀ = 1.71). Their stochastic SEIRD ensemble is calibrated by simulation filtering to 598 cumulative confirmed cases on 8 June 2026, with the reporting fraction fixed at 1.0. These are therefore projected confirmed cases, a floor on the true size, rather than the ascertainment-corrected cumulative cases that this package's C_T and McCabe et al. estimate. The 8/10 June row is their calibration anchor and the dates from 24 June on are forward projections. The bounds are 90% prediction intervals.

source
BVDOutbreakSize.CHAMLA_CONFIRMED_W12 Constant
julia
CHAMLA_CONFIRMED_W12

Chamla et al. (2026) week-12 (24 June 2026) cumulative confirmed-case projection under their three transmissibility scenarios, as (scenario, median, lower_90, upper_90) tuples (low R₀ = 1.42, central R₀ = 1.71, high R₀ = 2.08). Week 12 is the forward horizon closest to this analysis's current cut-off, so the scenario spread sits beside the observed confirmed count and our matched-date projection. The bounds are 90% prediction intervals. The same estimand caveat as CHAMLA_CONFIRMED_CENTRAL applies. These are confirmed cases, not ascertainment-corrected total cases.

source
BVDOutbreakSize.FROZEN_FIT Constant
julia
FROZEN_FIT

Value of a fit column identifying a row as a frozen-fit forecast, the joint model re-fit and evaluated at a past cut-off (forecast_frozen.csv). That asset labels its joint rows with this value and its single-stream rows with each fit's own id. score_release applies it to every row of an archive published before the column existed.

source
BVDOutbreakSize.ITURI_DAILY_TRAVEL Constant
julia
ITURI_DAILY_TRAVEL

Default prior mean for the daily outbound traveller volume from Ituri Province across seven points of entry.

source
BVDOutbreakSize.ITURI_DAILY_TRAVEL_SD Constant
julia
ITURI_DAILY_TRAVEL_SD

Default prior SD for the daily outbound traveller volume, covering point-of-entry-to-point-of-entry variation and reporting uncertainty in the underlying mobility survey.

source
BVDOutbreakSize.ITURI_POPULATION Constant
julia
ITURI_POPULATION

Source population for the Ituri Province (McCabe et al., Table 1).

source
BVDOutbreakSize.JOINT_FIT Constant
julia
JOINT_FIT

Value of a fit column identifying a row as the joint model's forecast or estimate. A FROZEN_FIT row is drawn in the same role, being the same model re-fit at a past cut-off rather than a separate fit.

source
BVDOutbreakSize.PROVINCE_CENTRES Constant
julia
PROVINCE_CENTRES

Representative point of each patch, in PROVINCE_NAMES order, as (latitude, longitude) in decimal degrees north and east. A patch holding one province takes its centre from PROVINCE_SOURCE_CENTRES; a pooled patch takes the mean of its members' centres weighted by PROVINCE_SOURCE_POPULATIONS, the populations the kernel reads.

source
BVDOutbreakSize.PROVINCE_DISTANCE_DECAY Constant
julia
PROVINCE_DISTANCE_DECAY

Exponent on the distance term of province_importation_kernel. One is the conventional gravity value. It is fixed rather than sampled because the importation intensity it scales is already weakly identified against the secondary provinces' seeds, so a second free parameter on the same term would not be determined by anything.

source
BVDOutbreakSize.PROVINCE_LABELS Constant
julia
PROVINCE_LABELS

Display names for the patches, in PROVINCE_NAMES order, for table rows and figure panels.

source
BVDOutbreakSize.PROVINCE_MEMBERS Constant
julia
PROVINCE_MEMBERS

Which source provinces each patch pools, keyed by PROVINCE_NAMES and valued in PROVINCE_SOURCE_NAMES. Every source province belongs to exactly one patch, so the patches partition the national totals and the composition likelihoods stay conditional on them.

source
BVDOutbreakSize.PROVINCE_NAMES Constant
julia
PROVINCE_NAMES

The patches of the meta-population model, in patch order. The first entry is the primary patch. It is the origin of the outbreak and the reference for the Uganda export propensities in province_export_pressure_model. The per-patch reproduction-number deviations in patch_rt_model sum to zero, so no patch is their reference.

Ituri, Nord-Kivu and Haut-Uele carry enough confirmed cases to be patches in their own right. The rest hold a few dozen between them and are pooled into other. Giving each of those its own reproduction number and ascertainment would sample dimensions nothing informs, and their estimates would be the deviation prior read back.

source
BVDOutbreakSize.PROVINCE_POPULATIONS Constant
julia
PROVINCE_POPULATIONS

Resident population of each patch, in PROVINCE_NAMES order, summed over the provinces it pools, from the 2019 INS figures in PROVINCE_SOURCE_POPULATIONS. Used to weight the between-province importation kernel, to centre the province background and bed-capacity shares, and as each patch's susceptible pool in the renewal.

source
BVDOutbreakSize.PROVINCE_SOURCE_CENTRES Constant
julia
PROVINCE_SOURCE_CENTRES

Population-weighted centre of each province in PROVINCE_SOURCE_NAMES order, as (latitude, longitude) in decimal degrees north and east, the points the distance term in province_importation_kernel reads. Each is the mean of the province's health-zone centroids weighted by the zones' WorldPop counts (WorldPop 2025, University of Southampton, https://www.worldpop.org, CC BY 4.0), with the zone boundaries from the Ministry of Health DRC_Health_zones shapefile on the Humanitarian Data Exchange (https://data.humdata.org/dataset/drc-health-data). Both come from the INRB-UMIE build of the health-zone map (https://github.com/INRB-UMIE/BDBV2026-Data, build/drc_health_zones.geojson at commit f7d3907, 28 September 2026), and scripts/build_province_centres.py (task province-centres) prints this literal from it.

A capital can sit far from where a province's people live. Goma is at the southern tip of Nord-Kivu, about 110 km further from Bunia than Nord-Kivu's population centre.

source
BVDOutbreakSize.PROVINCE_SOURCE_NAMES Constant
julia
PROVINCE_SOURCE_NAMES

Every province the situation reports carry a per-province row for, in the order they first appear. These key the per-province blocks of data/observations.toml. They are not the model's patches: see PROVINCE_NAMES for those and PROVINCE_MEMBERS for how the two relate.

source
BVDOutbreakSize.PROVINCE_SOURCE_POPULATIONS Constant
julia
PROVINCE_SOURCE_POPULATIONS

Resident population of each province in PROVINCE_SOURCE_NAMES order. 2019 figures from the Democratic Republic of the Congo's Institut National de la Statistique, Annuaire statistique RDC 2020 (March 2021), as tabulated at https://en.wikipedia.org/wiki/Provinces_of_the_Democratic_Republic_of_the_Congo (accessed 15 September 2026).

One source for all seven rather than the best figure for each, so consistency between provinces matters more than the accuracy of any one of them. The relative sizes enter the importation kernel and centre the province background share. The absolute sizes are the pools the renewal depletes (renewal_infections, patch_infections), where they only bound the outbreak.

source
BVDOutbreakSize.RENEWAL_START_LEAD Constant
julia
RENEWAL_START_LEAD

Days the renewal start sits after the genetic TMRCA day. The renewal start is where the analytic cryptic phase hands off to the recursion. Placing it a lead after the TMRCA rather than exactly on it leaves the observed span τ_obs = n − renewal_start strictly shorter than tmrca_days, so the genetic censored bound on the total age T = m·G + τ_obs stays informative and bounds the cryptic duration m·G from below. The lead also allows for the TMRCA's own molecular-clock uncertainty.

source
BVDOutbreakSize.REPORT_SCENARIOS Constant
julia
REPORT_SCENARIOS

Published point estimates of cumulative cases C_T from McCabe et al. (Imperial College London, 20 May 2026 update), as (label, value) tuples in the order they appear in Tables 1 and 2. These are the scenario means. The matching 95% confidence intervals are carried by REPORT_SCENARIOS_CI.

source
BVDOutbreakSize.REPORT_SCENARIOS_CI Constant
julia
REPORT_SCENARIOS_CI

Published McCabe et al. scenario estimates with their reported 95% confidence intervals, for three vintages. The 18 May 2026 report (McCabe and others, 2026), the 20 May 2026 update (McCabe and others, 2026) and the Lancet Infectious Diseases publication (McCabe et al., 2026), whose inputs are as of 27 May 2026. Each entry is a (date, label, mean, lower, upper) tuple, where date is that vintage's own cut-off date.

Method 1 (geographic spread from exported cases and travel volume) is unchanged between the two Imperial reports, so it is recorded once under the 20 May vintage. Method 2 (back-calculation from deaths) differs. The 18 May report used 88 deaths and CFR 24/30/40%, the 20 May update 131 deaths and CFR 26/33/40%. Confidence intervals are exact negative-binomial (Method 1) and Poisson likelihood-profile (Method 2), as reported in Tables 1 and 2 of each report.

The 27 May Lancet vintage uses 240 deaths and three Uganda imports, and varies the epidemic doubling time T_d (7/10/14 d) for both methods rather than the earlier geographic window w and onset-to-death τ. Its back-calculation fixes the onset-to-death gamma (mean 11.37 d, SD 5.41) and assumes 30% of deaths are attributable to Ebola. The paper swaps the method numbers, so its "method 1" is the back-calculation. This package keeps its own convention (M1 geographic spread with negative-binomial CIs, M2 back-calculation with Poisson CIs), confirmed by the paper's reported CI types.

source
BVDOutbreakSize.RT_INTERVENTION_RAMP Constant
julia
RT_INTERVENTION_RAMP

Time scale in days of the logistic ramp over which the outbreak-response intervention takes effect on R_t, centred on the breakpoint (sigmoid_ramp). Ramped rather than an instantaneous step, because the response damps transmission over weeks. Both the model and every reconstruct_rt caller read this constant, so a reconstruction cannot drift from the value the model fitted.

source
BVDOutbreakSize.RT_WALK_LEAD Constant
julia
RT_WALK_LEAD

Days before the first situation report (breakpoint) at which the reproduction-number random walk is allowed to start moving, rather than holding R_t flat at R0 until the report. A month lets the walk capture transmission dynamics before the outbreak is first reported, while staying floored at the renewal start so the walk never precedes the seeded trajectory.

source
BVDOutbreakSize.haversine_km Method
julia
haversine_km(
    a::Tuple{Real, Real},
    b::Tuple{Real, Real}
) -> Any
julia
haversine_km(a, b)

Great-circle distance in kilometres between two (latitude, longitude) points in decimal degrees, on a spherical Earth of radius 6371 km.

source
BVDOutbreakSize.province_distance_matrix Function
julia
province_distance_matrix() -> Matrix{Float64}
province_distance_matrix(centres::AbstractVector) -> Any
julia
province_distance_matrix(centres = PROVINCE_CENTRES)

Great-circle distances in kilometres between every pair of province centres, as a symmetric matrix with a zero diagonal. Built from PROVINCE_CENTRES by haversine_km.

source
BVDOutbreakSize.province_importation_kernel Function
julia
province_importation_kernel(; ...) -> Matrix{Float64}
province_importation_kernel(
    pops::AbstractVector;
    distances,
    decay
) -> Any
julia
province_importation_kernel(pops = PROVINCE_POPULATIONS; distances, decay)

Between-province importation kernel K, where K[p, q] is the relative rate of infectious travel from province q into province p. Diagonal is zero (no self-importation). The overall intensity is carried by the sampled ε in patch_infection_model, so only the relative structure matters here.

This is a gravity kernel. Travel from q to p scales with the destination population and falls with the distance between the two population centres,

with γ fixed at PROVINCE_DISTANCE_DECAY and d from province_distance_matrix. Pass distances as a zero matrix to recover the population-only kernel.

Each column is scaled so that its off-diagonal entries sum to 1 - N_q/N, the share of the country that is not q itself. The distance term then changes where a province's exported transmission lands without changing how much of it leaves. Every column's off-diagonal sum stays below one at any ε in [0, 1], which stops a province exporting more transmission than it generates.

There is no origin-destination or mobility data for this outbreak, so the kernel is a structural assumption rather than a measurement, and ε is weakly identified against the secondary-patch seeds, since both can raise a secondary province's early incidence. Treat the split between imported and locally-seeded infections as poorly determined even though their sum is not.

source
BVDOutbreakSize.OBSERVATION_STREAMS Constant

One entry per observation stream the manifest loader carries, in the order the report presents them. Each entry is a NamedTuple:

  • id: the canonical short identifier used across the package.

  • field: the field of a loaded observation set holding the stream's dated history.

  • label: the display title the report uses for the stream.

  • score_label: the label the forecast archive and the release scoring table use, or nothing for an unscored stream.

  • forecast_prefix: the stem of the stream's forecast columns (:cases for cases_cum and cases_new), or nothing for a stream that is not forecast.

  • stratum: the spatial level the stream is reported at, :national, :province or :zone.

A :province or :zone stream is a block of dated series rather than one series: a Dict keyed by province, or by province and then health zone, each holding the same (; days, counts) shape. Its vintage grid is shared across the block, so stream_last_date reads the block's last vintage over all of its series.

This is the single list of streams, so a stream is named once rather than once per consumer. stream_id resolves any of the four vocabularies back to id.

source
BVDOutbreakSize.STREAM_REPORTING_GRACE_DAYS Constant

Days a stream's last vintage may lag the cut-off and still count as reporting. One reporting week, which is also the forecast horizon, so a stream that skips a single situation report is not read as stopped.

source
BVDOutbreakSize.confirmed_break_correction Method
julia
confirmed_break_correction(
    obs,
    from_day::Real,
    to_day::Real;
    deaths
) -> Float64

Total harmonisation correction carried by a confirmed stream over the grid days (from_day, to_day], summing confirmed_break_steps over every listed break day in the window.

The window is half open on the left, so a break day on the origin belongs to the window before it. deaths selects the confirmed-death stream rather than confirmed cases. Returns zero when the window holds no break day.

source
BVDOutbreakSize.freeze_observations Method
julia
freeze_observations(
    cutoff_date::Union{Dates.Date, AbstractString};
    path,
    seeding_lead
) -> NamedTuple{(:n, :cutoff, :seeding, :exported_cases, :exports_deaths, :export_case_days, :export_death_days, :total_deaths, :reported_cases, :confirmed_cases, :confirmed_deaths, :tests_analysed, :reported_history, :confirmed_history, :confirmed_deaths_history, :deaths_history, :lab_history, :lab_daily_history, :suspected_daily_history, :suspected_daily_deaths_history, :isolation_history, :bed_capacity_history, :recovered_history, :recovered_cases, :treatment_admissions_history, :treatment_deaths_history, :treatment_ruleout_history, :treatment_absconded_history, :treatment_confirmed_incare_history, :treatment_suspect_incare_history, :occupancy_break_days, :confirmed_break_days, :confirmed_break_gross_cases, :confirmed_break_gross_deaths, :tests_received_history, :onset_curve_history, :onset_report_history, :province_confirmed_history, :province_death_history, :province_lab_daily_history, :province_isolation_history, :province_bed_capacity_history, :province_admissions_history, :zone_confirmed_history, :zone_death_history, :tmrca_days, :who_first_sitrep_days), <:Tuple{Int64, Any, Any, Any, Any, Any, Any, Any, Any, Any, Any, Any, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, Any, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, Any, Any, Any, Any, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, Union{@NamedTuple{onset_days::Vector{Int64}, report_days::Vector{Int64}, prev_report_days::Vector{Int64}, increments::Vector{Int64}, total_days::Vector{Int64}, total_counts::Vector{Int64}, last_total::Missing}, @NamedTuple{onset_days::Vector{Int64}, report_days::Vector{Int64}, prev_report_days::Vector{Int64}, increments::Vector{Int64}, total_days::Vector{Int64}, total_counts::Vector{Int64}, last_total::Int64}}, @NamedTuple{days::Vector{Int64}, counts::Vector{Int64}}, Dict{String, @NamedTuple{days::Vector{Int64}, counts::Vector{Int64}}}, Dict{String, @NamedTuple{days::Vector{Int64}, counts::Vector{Int64}}}, Dict{String, @NamedTuple{days::Vector{Int64}, counts::Vector{Int64}}}, Dict{String, @NamedTuple{days::Vector{Int64}, counts::Vector{Int64}}}, Dict{String, @NamedTuple{days::Vector{Int64}, counts::Vector{Int64}}}, Dict{String, @NamedTuple{days::Vector{Int64}, counts::Vector{Int64}}}, Dict{String, Dict{String, @NamedTuple{days::Vector{Int64}, counts::Vector{Int64}}}}, Dict{String, Dict{String, @NamedTuple{days::Vector{Int64}, counts::Vector{Int64}}}}, Int64, Any}}
julia
freeze_observations(cutoff_date; path = default manifest)

Load the observation manifest frozen to cutoff_date. Every dated history is truncated to the vintages available by then, so the returned named tuple is what the renewal model would have seen on that date. The cut-off scalar totals are taken from the truncated histories rather than the manifest's full-data scalars. Use it to re-evaluate the renewal estimate at a past report date for a matched-in-time comparison.

cutoff_date accepts a Date or an ISO date string. It must be on or after the earliest history vintage in the manifest, since the DRC series begins 18 May 2026. An earlier date leaves the suspected streams empty and is not a meaningful renewal fit.

source
BVDOutbreakSize.history_first_date Method
julia
history_first_date(
    grid_date,
    history
) -> Union{Missing, Dates.Date}

Date a dated history was first reported, or missing when it has no vintages. A history is any (; days) series of day-indices, and grid_date maps one of those indices to its calendar date, so a per-province or an assembled series reads the same way a manifest field does. Before the first vintage a cumulative total reads zero because the series has not started, which is the absence of a series rather than a count of zero.

source
BVDOutbreakSize.history_last_date Method
julia
history_last_date(
    grid_date,
    history
) -> Union{Missing, Dates.Date}

Date a dated history was last reported, or missing when it has no vintages. A history is any (; days) series of day-indices, and grid_date maps one of those indices to its calendar date. Past the last vintage the series is only ever repeated at its last reported value rather than genuinely observed.

source
BVDOutbreakSize.load_health_zones Function
julia
load_health_zones(

) -> Vector{@NamedTuple{zone::String, label::String, province::String, population::Int64, lat::Float64, lon::Float64, zscode::String}}
load_health_zones(
    path::AbstractString
) -> Vector{@NamedTuple{zone::String, label::String, province::String, population::Int64, lat::Float64, lon::Float64, zscode::String}}
julia
load_health_zones(path = data/health_zones.csv)

Read the health-zone metadata table: one named tuple per zone with the manifest key zone, the display label, the province key, the WorldPop population, the polygon centroid lat and lon in decimal degrees, and the DHIS2 zscode from the health-zone shapefile. Rows are in patch order then alphabetical by key. The file is written by scripts/build_health_zones.py; see data/README.md for its sources.

source
BVDOutbreakSize.load_observations Function
julia
load_observations(
;
    ...
) -> NamedTuple{(:n, :cutoff, :seeding, :exported_cases, :exports_deaths, :export_case_days, :export_death_days, :total_deaths, :reported_cases, :confirmed_cases, :confirmed_deaths, :tests_analysed, :reported_history, :confirmed_history, :confirmed_deaths_history, :deaths_history, :lab_history, :lab_daily_history, :suspected_daily_history, :suspected_daily_deaths_history, :isolation_history, :bed_capacity_history, :recovered_history, :recovered_cases, :treatment_admissions_history, :treatment_deaths_history, :treatment_ruleout_history, :treatment_absconded_history, :treatment_confirmed_incare_history, :treatment_suspect_incare_history, :occupancy_break_days, :confirmed_break_days, :confirmed_break_gross_cases, :confirmed_break_gross_deaths, :tests_received_history, :onset_curve_history, :onset_report_history, :province_confirmed_history, :province_death_history, :province_lab_daily_history, :province_isolation_history, :province_bed_capacity_history, :province_admissions_history, :zone_confirmed_history, :zone_death_history, :tmrca_days, :who_first_sitrep_days), <:Tuple{Int64, Any, Any, Any, Any, Any, Any, Any, Any, Any, Any, Any, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, Any, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, Any, Any, Any, Any, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, Union{@NamedTuple{onset_days::Vector{Int64}, report_days::Vector{Int64}, prev_report_days::Vector{Int64}, increments::Vector{Int64}, total_days::Vector{Int64}, total_counts::Vector{Int64}, last_total::Missing}, @NamedTuple{onset_days::Vector{Int64}, report_days::Vector{Int64}, prev_report_days::Vector{Int64}, increments::Vector{Int64}, total_days::Vector{Int64}, total_counts::Vector{Int64}, last_total::Int64}}, @NamedTuple{days::Vector{Int64}, counts::Vector{Int64}}, Dict{String, @NamedTuple{days::Vector{Int64}, counts::Vector{Int64}}}, Dict{String, @NamedTuple{days::Vector{Int64}, counts::Vector{Int64}}}, Dict{String, @NamedTuple{days::Vector{Int64}, counts::Vector{Int64}}}, Dict{String, @NamedTuple{days::Vector{Int64}, counts::Vector{Int64}}}, Dict{String, @NamedTuple{days::Vector{Int64}, counts::Vector{Int64}}}, Dict{String, @NamedTuple{days::Vector{Int64}, counts::Vector{Int64}}}, Dict{String, Dict{String, @NamedTuple{days::Vector{Int64}, counts::Vector{Int64}}}}, Dict{String, Dict{String, @NamedTuple{days::Vector{Int64}, counts::Vector{Int64}}}}, Int64, Any}}
load_observations(
    path::AbstractString;
    seeding_lead,
    cutoff_date,
    onset_curve_path
) -> NamedTuple{(:n, :cutoff, :seeding, :exported_cases, :exports_deaths, :export_case_days, :export_death_days, :total_deaths, :reported_cases, :confirmed_cases, :confirmed_deaths, :tests_analysed, :reported_history, :confirmed_history, :confirmed_deaths_history, :deaths_history, :lab_history, :lab_daily_history, :suspected_daily_history, :suspected_daily_deaths_history, :isolation_history, :bed_capacity_history, :recovered_history, :recovered_cases, :treatment_admissions_history, :treatment_deaths_history, :treatment_ruleout_history, :treatment_absconded_history, :treatment_confirmed_incare_history, :treatment_suspect_incare_history, :occupancy_break_days, :confirmed_break_days, :confirmed_break_gross_cases, :confirmed_break_gross_deaths, :tests_received_history, :onset_curve_history, :onset_report_history, :province_confirmed_history, :province_death_history, :province_lab_daily_history, :province_isolation_history, :province_bed_capacity_history, :province_admissions_history, :zone_confirmed_history, :zone_death_history, :tmrca_days, :who_first_sitrep_days), <:Tuple{Int64, Any, Any, Any, Any, Any, Any, Any, Any, Any, Any, Any, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, Any, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, Any, Any, Any, Any, NamedTuple{(:days, :counts), <:Tuple{Any, Any}}, Union{@NamedTuple{onset_days::Vector{Int64}, report_days::Vector{Int64}, prev_report_days::Vector{Int64}, increments::Vector{Int64}, total_days::Vector{Int64}, total_counts::Vector{Int64}, last_total::Missing}, @NamedTuple{onset_days::Vector{Int64}, report_days::Vector{Int64}, prev_report_days::Vector{Int64}, increments::Vector{Int64}, total_days::Vector{Int64}, total_counts::Vector{Int64}, last_total::Int64}}, @NamedTuple{days::Vector{Int64}, counts::Vector{Int64}}, Dict{String, @NamedTuple{days::Vector{Int64}, counts::Vector{Int64}}}, Dict{String, @NamedTuple{days::Vector{Int64}, counts::Vector{Int64}}}, Dict{String, @NamedTuple{days::Vector{Int64}, counts::Vector{Int64}}}, Dict{String, @NamedTuple{days::Vector{Int64}, counts::Vector{Int64}}}, Dict{String, @NamedTuple{days::Vector{Int64}, counts::Vector{Int64}}}, Dict{String, @NamedTuple{days::Vector{Int64}, counts::Vector{Int64}}}, Dict{String, Dict{String, @NamedTuple{days::Vector{Int64}, counts::Vector{Int64}}}}, Dict{String, Dict{String, @NamedTuple{days::Vector{Int64}, counts::Vector{Int64}}}}, Int64, Any}}

Load the BVD observation manifest from path (a dated TOML file) and return a named tuple for the renewal model. Calendar dates are converted to 1-based grid day-indices, day 1 the seeding day and day n the cut-off, using the cut-off (as_of_date) and a seeding day placed seeding_lead days before the genetic TMRCA date.

Returns the grid length n, the cutoff and seeding dates, and the per-stream cumulative totals at the cut-off (reported_cases, total_deaths, confirmed_cases, confirmed_deaths, tests_analysed, exported_cases, exports_deaths). A scalar with no explicit TOML block is taken from the final vintage of the matching history.

The dated Uganda export series are grid day-indices (export_case_days, export_death_days), each a sorted list of detection or death days on or before the cut-off.

The per-vintage histories are returned as (; days, counts) with days the grid day-indices: reported_history, confirmed_history, confirmed_deaths_history, deaths_history, lab_history (cumulative analysed specimens), lab_daily_history (24h analysed counts), suspected_daily_history (daily new-suspect inflow), suspected_daily_deaths_history (daily new suspected deaths), isolation_history (daily isolation occupancy), bed_capacity_history (occupancy / reported occupancy rate), recovered_history (cumulative recovered among confirmed), treatment_confirmed_incare_history and treatment_suspect_incare_history (the occupancy split into two prevalence sub-stocks that sum to the total), and tests_received_history.

The digitised symptom-onset reporting triangle is returned as onset_curve_history, the per-vintage increments read from onset_curve_path, by default onset_curve_scanned.csv alongside path. A manifest read from elsewhere has no such sibling, so name the triangle explicitly to keep the stream rather than degrade it to a no-op. See load_onset_curve for the dedup and increment construction. The same triangle's cumulative confirmed-by-onset total is returned as onset_report_history in the usual shape.

Also returned are the genetic TMRCA bound tmrca_days (days before the cut-off) and who_first_sitrep_days (days from the earliest reported-case vintage to the cut-off). The intervention breakpoint grid day is n - who_first_sitrep_days.

source
BVDOutbreakSize.province_care_observations Function
julia
province_care_observations(
    history;
    ...
) -> NamedTuple{(:days, :patches, :counts), <:Tuple{Vector{Int64}, Any, Any}}
province_care_observations(
    history,
    province_names::AbstractVector;
    members,
    every,
    changes_only
) -> NamedTuple{(:days, :patches, :counts), <:Tuple{Vector{Int64}, Any, Any}}
julia
province_care_observations(history, province_names; members, every,
                           changes_only)

Long-format province observations for the isolation-occupancy and bed splits in treatment_flow_model: one row per (day, patch) with a printed figure, (; days, patches, counts) sorted by day. history maps a source province to its sparse (days, counts) series, as load_observations reads the province_isolation_history and province_bed_capacity_history blocks. A pooled patch (see PROVINCE_MEMBERS) is present on a day only when every member that has printed on or before that day prints that day, and its count is their sum, since a partial sum would read a silent member as an empty ward. A province absent from history never contributes.

The split scores each kept day as a fresh draw, and a stock reprinted daily is not one: the same in-patients fill the same beds from one report to the next. every = 7 keeps one day in seven of the days with a split, about one bed stay apart, so the kept days are close to independent; with every day kept the split over-constrains the per-patch demand and slows the sampler. changes_only = true keeps a day only when some province's count differs from its last kept count, for the bed counts, which the reports repeat unchanged for weeks.

source
BVDOutbreakSize.province_increment_matrix Method
julia
province_increment_matrix(
    province_history,
    province_names::AbstractVector,
    n_patches::Integer
) -> NamedTuple{(:days, :increments), <:Tuple{Any, Matrix{Int64}}}
julia
province_increment_matrix(province_history, province_names, n_patches)

Reshape the per-province cumulative histories loaded by load_observations into the (n_patches × n_vintages) matrix of new-confirmed counts that province_composition_model scores, together with the shared vintage day indices.

Every province must be reported on the same vintage days, which the composition likelihood requires. It allocates each vintage's national total across the provinces, so a province missing from a vintage would silently shift cases into the others. A mismatch is an error.

Returns (; days, increments). When no per-province data is supplied, days is empty and the caller skips the composition term.

source
BVDOutbreakSize.province_lab_increment_matrix Function
julia
province_lab_increment_matrix(
    province_lab_daily_history;
    ...
) -> NamedTuple{(:days, :bins, :increments), <:Tuple{Any, Any, Any}}
province_lab_increment_matrix(
    province_lab_daily_history,
    province_names::AbstractVector;
    ...
) -> NamedTuple{(:days, :bins, :increments), <:Tuple{Any, Any, Any}}
province_lab_increment_matrix(
    province_lab_daily_history,
    province_names::AbstractVector,
    n_patches::Integer;
    every
) -> NamedTuple{(:days, :bins, :increments), <:Tuple{Any, Any, Any}}
julia
province_lab_increment_matrix(province_lab_daily_history, province_names,
                              n_patches; every)

Reshape the per-province daily analysed-specimen histories loaded by load_observations into the (n_patches × n_bins) matrix of analysed counts that the laboratory composition in bvd_joint scores, with the printed day indices days and the bin each day falls in (bins). every = 1 scores each printed day as its own bin. every = 7 sums the printed days into calendar weeks from the first day, so the composition scores a weekly split. A pooled patch sums its members' counts (see PROVINCE_MEMBERS). Every province must be reported on the same days, which the composition requires. Returns empty days when the history is absent or a patch has no analysed series, and the caller skips the term.

source
BVDOutbreakSize.stream_first_date Method
julia
stream_first_date(obs, stream) -> Union{Missing, Dates.Date}

Date stream was first reported in obs, or missing when obs does not carry the stream or the stream has no vintages. This is the date of the stream's first vintage, where its series begins.

The Uganda exports are a dated list of detections rather than a series of vintages, so their first reported date is the earlier of the first detected import and the first detected import death.

A :province or :zone stream is a block of dated series, so its first reported date is the vintage its block was first printed on.

source
BVDOutbreakSize.stream_forecast_columns Method
julia
stream_forecast_columns(
    stream
) -> Union{Nothing, @NamedTuple{cum::Symbol, new::Symbol}}

Forecast column names for stream as (; cum, new), or nothing for a stream the forecast does not carry. Built from the registry's forecast_prefix, so the column names are derived rather than repeated.

source
BVDOutbreakSize.stream_id Method
julia
stream_id(stream) -> Symbol

Resolve stream to its canonical identifier in OBSERVATION_STREAMS. Accepts a canonical identifier, an observation-set history field name, a scoring label, or a forecast column name (:cases_cum and :cases_new both resolve to :suspected_cases). The four vocabularies are disjoint, so the mapping is unambiguous. Errors on an unknown stream, naming the identifiers it knows.

source
BVDOutbreakSize.stream_last_date Method
julia
stream_last_date(obs, stream) -> Union{Missing, Dates.Date}

Date stream was last reported in obs, or missing when obs does not carry the stream or the stream has no vintages. This is the date of the stream's last vintage, since past it the series is only ever repeated at its last reported value rather than genuinely observed.

The Uganda exports are a dated list of detections rather than a series of vintages, so their last reported date is the later of the last detected import and the last detected import death.

A :province or :zone stream is a block of dated series, so its last reported date is the last vintage over all of them. A block the manifest carries as empty reports missing.

source
BVDOutbreakSize.stream_report_status Method
julia
stream_report_status(
    obs;
    grace,
    stratum
) -> DataFrames.DataFrame

Reporting status of every stream obs carries, one row per entry of OBSERVATION_STREAMS and in that order. Columns: the canonical stream identifier, its display label, the last_date it reported, whether it is still reporting at the cut-off, and days_since that last report.

stratum keeps the streams of one spatial level, :national, :province or :zone. Each report page reads the currency of its own level.

source
BVDOutbreakSize.stream_reporting Method
julia
stream_reporting(obs, stream; grace) -> Bool

Whether stream was still being reported at the cut-off of obs, that is whether its last vintage falls within grace days of the cut-off. A stream obs does not carry, or one with no vintages, is not reporting.

source
BVDOutbreakSize.zone_cumulative_falls Method
julia
zone_cumulative_falls(
    zone_history;
    min_fall
) -> Vector{@NamedTuple{province::String, zone::String, day::Int64, from::Int64, to::Int64}}
julia
zone_cumulative_falls(zone_history; min_fall = 1)

Every place a named zone's cumulative count falls between consecutive vintages by more than min_fall, as a vector of named tuples (; province, zone, day, from, to). A cumulative series cannot fall on its own, so each entry is a revision: the report has moved counts between zones, provinces or the unallocated row. min_fall guards against a single-unit correction being read as a revision. zone_increment_matrix clamps these falls to zero.

source
BVDOutbreakSize.zone_increment_matrix Function
julia
zone_increment_matrix(
    zone_history,
    patch_names::AbstractVector;
    ...
) -> Vector{@NamedTuple{patch::String, zones::Vector{Tuple{String, String}}, days::Vector{Int64}, increments::Matrix{Int64}, totals::Vector{Int64}, excluded::Vector{Int64}}}
zone_increment_matrix(
    zone_history,
    patch_names::AbstractVector,
    members::AbstractDict;
    reattribution
) -> Vector{@NamedTuple{patch::String, zones::Vector{Tuple{String, String}}, days::Vector{Int64}, increments::Matrix{Int64}, totals::Vector{Int64}, excluded::Vector{Int64}}}
julia
zone_increment_matrix(zone_history, patch_names, members = PROVINCE_MEMBERS;
    reattribution = zone_reattribution_days(zone_history))

Reshape the per-health-zone cumulative histories loaded by load_observations into one increment matrix per patch, the within-patch analogue of province_increment_matrix.

Each patch pools the source provinces members gives it (a name with no entry is its own province), and its matrix has one row per zone of those provinces, the unallocated rows left out, and one column per vintage. Every zone must be reported on the same vintage days, which the tables guarantee by sharing one dates array; a mismatch is an error. Increments are the differences of consecutive cumulative counts, the first from zero, clamped at zero as for the provinces.

A vintage on which a member province's unallocated count falls (reattribution, the days per province of zone_reattribution_days) is a reattribution of counts the report had carried as unallocated into named zones. The zone increments would read them as new cases, so that column is set to zero for the patch, which drops the cell from the composition, and its days are returned as excluded. The next vintage's increment is still the difference of the cumulative counts, so nothing is counted twice. Pass the merged days of more than one history (the death block's falls as well as the case block's) to exclude the union, or an empty Dict to keep every vintage.

Returns a vector with one named tuple per patch: patch (its name), zones (a vector of (province, zone) key pairs in row order), days, increments (the (n_zones × n_vintages) matrix), totals (the column sums, the allocated count each vintage's composition conditions on) and excluded (the zeroed vintage days). A patch none of whose provinces has zone data gets an empty matrix. An empty zone_history returns an empty vector.

source
BVDOutbreakSize.zone_reattribution_days Method
julia
zone_reattribution_days(
    zone_history;
    include_zone_falls,
    min_fall
) -> Dict{String, Vector{Int64}}
julia
zone_reattribution_days(zone_history; include_zone_falls = false,
    min_fall = 1)

The vintage days on which a province's unallocated cumulative count falls, keyed by province, from the per-health-zone histories loaded by load_observations. A fall means the report has attributed cases (or deaths) it had carried as unallocated to named zones, so on that day the zones' cumulative counts rise by more than the province's. A province whose unallocated row never falls, or that has none, is absent. zone_increment_matrix leaves those vintages out of the composition.

A revision can also move counts the other way, out of named zones, leaving the unallocated row flat or rising. That vintage is a revision just as much, but the unallocated rule does not see it. With include_zone_falls the vintages of zone_cumulative_falls are added, so any vintage on which a named zone loses more than min_fall is left out too. It is off by default.

source
BVDOutbreakSize.ONSET_REPORT_MAX_DELAY Constant
julia
ONSET_REPORT_MAX_DELAY

Maximum symptom-onset-to-report delay (days) the hazard model tracks, d = 0 … D-1. The digitised triangle's own between-vintage increments settle to noise by about three weeks. Roughly 90% of a bar's eventual total is in by delay 14-17 and 95-97% by 20-25, and the tail beyond that is scan noise. 28 sits where the reporting signal has decayed into that noise floor.

source
BVDOutbreakSize.load_onset_curve Method
julia
load_onset_curve(path::AbstractString; cutoff, seeding)
julia
load_onset_curve(path; cutoff, seeding)

Load the digitised symptom-onset reporting triangle at path and build the cells onset_reporting_model fits: each onset date's level at its first print, then its corrections while the reporting delay still moves it.

Distinct SitRep vintages are recovered by exact-value dedup (_dedup_onset_blocks), so reprinted figures collapse to their earliest report date. Vintages reported after cutoff are dropped, the same convention load_observations uses for every other stream, so advancing the manifest as_of_date past a newly-digitised vintage's report date picks that vintage up with no code change.

Every onset date is scored once as a level and then only through increments, so no printed count enters the likelihood twice:

  • Level. The first vintage that prints onset date u gives one cell, differenced against an implicit empty predecessor (the sentinel prev_report_days[i] = 0), right-truncated at that vintage's report day. This is the only cell for dates first printed past the delay support, which is the complete curve back to the start of the digitised window.

  • Corrections. Each later vintage s that prints u while report_day(s) - u < ONSET_REPORT_MAX_DELAY gives one cell against the last earlier vintage that printed u:

julia
y = confirmed_total(s, u) - confirmed_total(prev, u)

A level plus its corrections telescopes to the latest print inside the delay support. Past the support the modelled increment is zero, so later reprints of a settled date are not scored and late reclassification is not modelled.

Each block has its own printed extent, its earliest to latest digitised onset date, and the published figures stop their x axis short of the report date by anything from zero to eight days. The axis simply ends, with substantial counts on the last printed bar, rather than running on with zero-height bars. A date outside a vintage's extent is unobserved in that vintage rather than zero, so it gets no cell there, and its next correction is taken against the last vintage that did print it. The choice of axis limit does not depend on the counts it hides, so this treats them as missing at random. Inside a block's own extent a missing row is a digitisation omission of a zero-height bar and does read as a true zero.

onset_report_cdf returns 0 for any negative delay and grid day 0 postdates no valid onset day, so the sentinel recovers the level as a difference from nothing, with no extra branch downstream. Level cells are what pin the ascertainment. Corrections only pin differences of F, so without a level somewhere the ascertainment would float.

Returns (; onset_days, report_days, prev_report_days, increments, total_days, total_counts, last_total). The first four are length-matched Vector{Int}s (1-based grid day-indices for the first three, the observed cell for the fourth) ready for onset_reporting_model. total_days and total_counts are the cumulative confirmed total printed by each surviving vintage, keyed on its report day, in the same (days, counts) shape every other stream's history carries. They are built from every printed bar of a vintage, not from the scored cells. last_total is the final entry of total_counts, or missing when no vintage survives.

The per-vintage totals are not monotone across vintages. Scan error on each figure means a later scan can read a smaller total than an earlier one even though late reporting only ever adds cases, and this happens more than once in the current data. Consumers that need a non-decreasing series must say what they do with a fall rather than assume it cannot happen, and anything scored against this series should be scored on its increments rather than its level.

A missing path, or a manifest with no in-cutoff vintage, returns the same empty, missing-total shape, so the stream degrades to a no-op rather than throwing.

source