CensoBR.jl is a Julia package for downloading and processing microdata from the Brazilian Population Census (Censo Demográfico).
In specific, this package:
- Downloads census data directly from the IBGE servers
- Process these data using the IBGE's data dictionary
- Saves the processed data as a
.parquetfile at a pre-selectedcachedirfolder - Further calls of using the same
cachedirfolder reuse the same processed data
CensoBR.jl currently supports the 2000 and 2010 Population Censuses.
This project is heavily inspired by the R package {censobr}.
Note: CensoBR.jl is under active development. The public API may change before the first stable release.
The main entry point is fetchcensus, which downloads, processes, and caches Census microdata. By default, CensoBR.jl uses the operating system's standard cache directory. A different directory can be specified with the cachedir keyword argument.
using CensoBR
dspath = fetchcensus(2000, :rj, :household)The returned object is a String indicating the resulting .parquet file path.
During the first call to fetchcensus for a given UF-year pair, the function takes advantage that the downloaded .zip archive contains data for all available record types and generates the corresponding .parquet files before cleaning up the raw data. This may make the initial call slower, but avoids repeated downloads and processing in subsequent calls for other record types.
To build a country-level file by stacking all states and the Federal District, pass :br as the UF:
dspath = fetchcensus(2010, :br, :person)This downloads and processes each UF in sequence. However, it does not generate intermediate .parquet files for individual UFs and therefore should not be used as a way to compile the complete set of UF sources.
Because CensoBR.jl caches processed data as .parquet, the files can be queried directly with tools such as DuckDB.jl, without first loading the entire Census dataset into memory.
using CensoBR
using DuckDB, DBInterface, DataFrames
dspath = fetchcensus(2000, :rj, :household; cachedir="path/to/cache")
con = DBInterface.connect(DuckDB.DB())
results = DBInterface.execute(
con,
"""
SELECT V1001, V0211, M0213
FROM read_parquet('$parquetpath')
WHERE V0211 = '4'
"""
)
df = DataFrame(results)Here, DuckDB performs the filtering and column selection directly against the .parquet file. Only the query result is materialized as a DataFrame.
CensoBR.jl includes variable metadata derived from the official IBGE documentation. Metadata can be accessed with fieldlabel, fieldvalues, and fieldnotes, or retrieved together with fieldmetadata.
For example:
julia> fieldmetadata(2000, :household, :V0211)Label:
TIPO DE ESCOADOURO
Values:
1 ⇒ Rede geral de esgoto ou pluvial
2 ⇒ Fossa séptica
3 ⇒ Fossa rudimentar
4 ⇒ Vala
5 ⇒ Rio, lago ou mar
6 ⇒ Outro escoadouro
Branco ⇒ para domicílio particular improvisado, domicílio coletivo e domicílio particular permanente que tinha banheiro(s) ou sanitário
Notes:
—
Individual components can be retrieved with:
# specific functions
fieldlabel(2000, :household, :V0211)
fieldvalues(2000, :household, :V0211)
fieldnotes(2000, :household, :V0211)
# using the Struct
metadata = fieldmetadata(2000, :household, :V0211)
metadata.label
metadata.values
metadata.notesMetadata access does not require downloading the Census microdata.
| Census | Record | Example |
|---|---|---|
| 2000 | Household | fetchcensus(2000, :rj, :household) |
| 2000 | Family | fetchcensus(2000, :rj, :family) |
| 2000 | Person | fetchcensus(2000, :rj, :person) |
| 2010 | Household | fetchcensus(2010, :rj, :household) |
| 2010 | Person | fetchcensus(2010, :rj, :person) |
| 2010 | Emigration | fetchcensus(2010, :rj, :emigration) |
| 2010 | Mortality | fetchcensus(2010, :rj, :mortality) |
| 2000, 2010 | Any supported record | fetchcensus(year, :br, record) |
CensoBR.jl supports Census microdata for all Brazilian states and the Federal District, and can stack them into a national file with :br.
IBGE Census microdata are distributed as fixed-width text files. CensoBR.jl uses bundled layouts derived from official IBGE documentation to determine each variable's byte position, width, type, and implied decimal places.
The parser follows these conventions:
- blank fields become
missing; - character fields become
String; - integer numeric fields become
Int; - numeric fields with implied decimal places become
Float64; and - character identifiers retain leading zeroes.
Parsing is performed internally as part of the conversion pipeline.