Global Justice Lab · Policing and Community Safety Research project ·
Munk School of Global Affairs & Public Policy, University of Toronto · Ron Levi ·
compiled with Claude Code (Anthropic)
Methods: where every number on this site comes from
Part of the Indigenous Survey Questions Explorer:
all 465 questions Canadian pollsters and researchers have asked about Indigenous issues since 1995.
Catalogue · evidence map ·
justice chart · health chart ·
attitude space (MCA) · meta-analysis.
1. The default rule: quoted, not estimated
Every figure in the catalogue is transcribed from a primary source fetched during compilation
(a report, data-table PDF, banner table, or official release), with the source URL, the survey base and
sample size, and any collapsing noted in the entry itself. Nothing is estimated, and a figure from one
population is never silently substituted for another. Where a source could not be found, the entry says
so, with the date searched.
2. Derived arithmetically from published figures
Some displayed numbers are simple arithmetic on published figures. Each carries a basis flag in
the harmonized dataset and a note in
its entry:
- Sums of published components (for example, agree = the published strongly-agree plus
somewhat-agree; food insecurity 50.8 = the published 37.7 moderate + 13.1 severe).
- Published nets, used exactly as the source printed them.
- Residuals (100 minus the published opposing share), used only where a release printed one
side of a collapsed binary; flagged
residual because an unreported don't-know share would
shrink the true figure. Five Ipsos rows carry this flag.
- The CRB index transform: the Reconciliation Barometer publishes means on a -1..+1 scale, not
distributions; we rescale to 0-100 as percent of scale range, and never pool these with true percentages.
3. Computed from public microdata (labelled in every entry)
Two groups of entries go beyond quotation, both authorized as a deliberate exception on 2026-08-04
and both fully reproducible from scripts in the
project repository:
- Canadian Election Study (23 entries). The CES releases microdata and codebooks, not toplines,
so we computed weighted distributions ourselves from the public files (Harvard Dataverse and Borealis),
with the weight, base and don't-know handling stated in each entry and the label COMPUTED, NOT
PUBLISHED. Script:
compute_ces_toplines.py. CES online samples are non-probability panels
weighted to census margins.
- StatCan public use microdata files (the PUMF pass). The APS 2017 and APS 2012, GSS-34 (2019)
and CHS 2018 PUMFs are public files, downloaded from UBC's Abacus Data Network under the Statistics
Canada Open Licence. Two situations arise. For the APS files, the PUMF data dictionaries themselves
publish the full weighted distributions, so those entries cite a published document and our
computation merely verifies it (it matches to the decimal on every check). For the GSS courts-confidence
and CHS housing-arrears items, no publication reports the by-Indigenous-identity breakdowns, so those
are our computations, labelled COMPUTED, NOT PUBLISHED, with the caveats stated (the GSS identity
question was asked only of Canada- and US-born respondents; the CHS measures the presence of an
Indigenous household member as reported by the reference person and does not aim to enumerate the
Indigenous population). Script:
compute_pumf_toplines.py. All are point estimates on the
main survey weights, without bootstrap confidence intervals.
- Two further StatCan public files (added 2026-08-05). GSS-35 (Social Identity, 2020) carries a
confidence-in-police item, CII_10, and the crowdsourced Impacts of COVID-19 on Canadians: Experiences of
Discrimination carries a trust-in-police item, TII_05A. Neither by-identity breakdown is published, so both
are labelled COMPUTED, NOT PUBLISHED. For GSS-35 our weighted all-respondent distribution reproduces the
published data-dictionary figures exactly. The COVID discrimination file additionally carries a
DO NOT POOL flag: see section 6. Script:
compute_pumf_toplines.py.
- The Residential School Denialism Survey (added 2026-08-05). The nine denialism items are
published in Figure 3 of Beauvais and Williamson (2025), and our computation from the public CC0
replication file reproduces all sixty-three printed cells, so those entries cite a published source that we
merely verify. The seven Indigenous resentment items and the three knowledge-quiz items have no published
per-item figure, so those are COMPUTED, NOT PUBLISHED. Every figure is for the untreated control group only
(n=960), unweighted, with don't-know in the denominator, matching the article's own base. Script:
compute_rsd_toplines.py.
4. What we do not have, and why
As of August 2026, 424 of the 465 items carry their own figures. The remaining 41 lack their own
published topline; their entries say exactly what exists instead:
| Items | Situation |
| 32 | The entry carries the nearest published figure with the difference stated: a
composite where per-item numbers were never released (the RHS food-security statements and K10 distress
items), a caregiver-reported figure where the youth topline is unpublished (FNREEES school items), or an
adjacent cycle or wording (several APS and CADS items). |
| 3 | EKOS items with no public release; only the firm's files could resolve them. |
| 4 | UAPS wordings that appear on neither the questionnaires nor the public data tables
(never fielded as catalogued; audits of 2026-08-05). |
| 1 | An APS 2012 aspiration item on StatCan's master file only (the bullying and racism
items gained published youth-report figures from the research literature). |
| 1 | A CES module whose variables could not be identified in the public file. |
The FNIGC boundary. The First Nations Regional Health Survey and the FNREEES are conducted by
the First Nations Information Governance Centre, and their microdata is governed under OCAP-based
principles: FNIGC holds it, and access runs through an application process by design. Everything FNIGC
published, we quote; everything derivable by arithmetic from those publications, we derived. The
remaining per-item gaps exist in no public object at all, so there is nothing to compute from; filling
them would require an FNIGC data-access application, not a calculation. We do not use FNIGC-held
microdata.
5. Reproducing it
The catalogue and every analysis rebuild from scripts in the
repository:
build_indigenous_survey_dataset_v2.py (the catalogue itself),
build_analysis_dataset.py, build_harmonized_outcome.py,
run_meta_analyses.py, robustness_gradient.py, run_ces_mca.py,
compute_ces_toplines.py, compute_pumf_toplines.py and
compute_rsd_toplines.py, plus the page builders.
The microdata sources are cited with DOIs and handles inside the scripts and entries.
6. The asterisk: sources that are not polling firms
An entry whose notes begin with * comes from a source that is not a polling firm. The marker was
introduced with the additions of 2026-08-05 and is applied from that date on; several long-standing sources
here are also not polling firms, among them Statistics Canada, the Canadian Election Study, the Department of
Justice and the First Nations Information Governance Centre. Three sources carry it now.
- The Residential School Denialism Survey (Beauvais and Williamson, Simon Fraser University and
McGill), 19 items. An academic survey fielded by Leger Opinion in November 2023 for a single peer-reviewed
article, not a public-opinion release. It is here because it meets the same sourcing test as the poll
entries and clears it on both halves: the wordings are published verbatim in the article and its
supplementary materials, the response distributions are published in the article's Figure 3, and the
microdata is public under CC0, so the figures can be checked rather than taken on trust. Two features of the
design are stated in every entry. Indigenous respondents were screened out of the sample by design, so these
are settler-opinion figures, as with the UAPS Non-Aboriginal Survey already catalogued here. And half the
sample first read a 250-word informational text about residential schools, so all figures are for the
untreated control group. One caveat we found and could not resolve: the deposited replication code
reverse-codes the "may not even contain Indigenous people" item inconsistently between its two scripts, a
difference of about two points on the article's pooled prevalence figure. We report the published per-item
figures, which the inconsistency does not touch.
- GSS-35 (Social Identity, 2020), one item. A Statistics Canada household survey, on the same basis
as the other Statistics Canada entries here.
- Impacts of COVID-19 on Canadians: Experiences of Discrimination (2020), one item, and the weakest
sample in the catalogue. A Statistics Canada crowdsourcing exercise for which, in its own words, "No sample
selection is done". It is here for the question wording; its distribution carries a DO NOT POOL flag
and is kept out of the harmonized supportive-share table, the trend series and the mood index, in the same
way as the Reconciliation Barometer index means. Statistics Canada advises that results "pertain only to the
participants and should not be used to draw conclusions about the general population" and "should not be
used to make comparisons with other probabilistic surveys". The selection bias also runs against this
particular measure in a known direction: Statistics Canada notes that people experiencing discrimination
were more inclined to participate, a skew it confirmed against the GSS on Victimization, so a trust-in-police
figure from this file is likely depressed by an unknown amount. Note too that the item asks about trust,
where the rest of the justice entries here ask about confidence.
Built 2026-08-05 · Global Justice Lab, Munk School of Global Affairs & Public
Policy, University of Toronto · Ron Levi · compiled, coded and analyzed with Claude Code
(Anthropic); figures transcribed from the cited primary sources; coding conventions reviewed and
accepted by Ron Levi.