geeViz.fsInsights.fia

FIADB-API client — Forest Inventory and Analysis estimates.

FIA is a probability sample of forest plots, not a census. Every estimate it produces is a design-based estimate with a sampling error, and this client is built around refusing to let you forget that.

A real query illustrates why. Forest area by county and forest type group for Alabama returns, among 4,241 plots:

Forest type group

Acres

SE (%)

Plots

Total

22,963,960

0.51

4241

Longleaf / slash pine group

1,184,334

6.07

258

White / red / jack pine group

15,748

54.92

4

That last row is a number with a 54.9% standard error resting on four plots. Rendered as a bare value in a chart it reads as fact. So estimate() returns se_pct and plots on every row, and flags the ones too thin to report.

Why plot count as well as standard error, when SE already grows as the sample shrinks: there are three regimes where SE cannot police itself. At n = 0 it is zero or undefined — reading as maximum precision when it means no information, which is the dangerous inversion. At n = 1 no variance estimate exists at all. And at small n the SE is itself estimated from that same tiny sample, so a small value can be luck rather than precision. There is a fourth, quieter problem: an SE is normally consumed as estimate ± 1.96·SE, which assumes approximate normality — at small n that interval undercovers, in the direction of overconfidence.

See Bechtold & Patterson, The Enhanced Forest Inventory and Analysis Program (GTR-SRS-80), for the estimation design itself.

Module Attributes

DEFAULT_MAX_SE_PCT

Flag a cell whose standard error exceeds this percentage.

DEFAULT_MIN_PLOTS

Flag a cell resting on fewer than this many plots.

FOREST_DEFINITIONS

Forest land definitions.

Functions

estimate(wc, snum, *[, rselected, ...])

Run an FIA estimate and return it tidy, with its sampling error.

reliable(df)

Drop cells flagged unreliable.

validate(wc, snum)

Check an attribute is answerable by an evaluation.

Exceptions

FIAValidationError

A request that would fail upstream, caught locally first.

geeViz.fsInsights.fia.DEFAULT_MAX_SE_PCT = 30.0

Flag a cell whose standard error exceeds this percentage.

geeViz.fsInsights.fia.DEFAULT_MIN_PLOTS = 30

Flag a cell resting on fewer than this many plots. 30 is the conventional floor for the normal approximation that turns an SE into a confidence interval — which is exactly the assumption ± 1.96·SE relies on.

geeViz.fsInsights.fia.FOREST_DEFINITIONS = ('FIADEF', 'RPADEF')

Forest land definitions. FIA and RPA produce different areas, and the API silently defaults to RPADEF — so a caller who never thinks about it gets numbers that quietly disagree with someone else’s EVALIDator pull.

exception geeViz.fsInsights.fia.FIAValidationError[source]

Bases: ValueError

A request that would fail upstream, caught locally first.

Every check here is cheap and offline. Turning an opaque server error into a specific local one matters most for agents, where a rejected call costs a whole turn.

geeViz.fsInsights.fia.validate(wc: int, snum: int) → None[source]

Check an attribute is answerable by an evaluation. Raises if not.

Attributes declare the evaluation type they need (EXPCURR, EXPVOL, EXPGROW, EXPMORT, EXPREMV, EXPCHNG, EXPDWM); evaluations advertise whether they support growth accounting. The pairing is knowable before the request leaves.

geeViz.fsInsights.fia.estimate(wc: int, snum: int, *, rselected: str = '', cselected: str = '', pselected: str = '', sdenom: int | None = None, forest_definition: str = 'FIADEF', str_filter: str = '', max_se_pct: float = 30.0, min_plots: int = 30, validate_first: bool = True) → Any[source]

Run an FIA estimate and return it tidy, with its sampling error.

Parameters:
  • wc – Evaluation group — see find_evaluations().

  • snum – Estimate attribute — see find_attributes().

  • rselected – Row grouping, as the exact display string from find_groupings().

  • cselected – Column grouping. Optional.

  • pselected – Page grouping. Optional.

  • sdenom – Denominator attribute, to produce a ratio estimate.

  • forest_definition –

    "FIADEF" or "RPADEF". Defaults to FIADEF and is always sent explicitly — the API’s own default is RPADEF, and leaving it implicit is how two people pull “the same” number and disagree.

    Caveat, unresolved. In testing, sending FIAorRPA=FIADEF still produced the echo “RPADEF as the forest land definition.” — so the parameter may be ignored, or may need a different spelling than the documentation gives. That is why the returned forest_definition column carries the API’s echo rather than what was requested: whatever the server actually applied is what a saved frame should record. Compare the two before publishing a number that depends on the distinction.

  • str_filter – SQL-style filter passed through as strFilter.

  • max_se_pct – Flag cells whose standard error exceeds this.

  • min_plots – Flag cells resting on fewer plots than this.

  • validate_first – Check attribute/evaluation compatibility locally before sending. Turn off only to probe the API directly.

Returns:

row, column, estimate, se_pct, plots, units, unreliable, unreliable_reason, plus attribute, evaluation and forest_definition for provenance.

Return type:

pandas.DataFrame with one row per cell

Raises:
  • FIAValidationError – The request would fail upstream.

  • UpstreamError / UpstreamUnavailable – The API refused or could not be reached.

geeViz.fsInsights.fia.reliable(df) → Any[source]

Drop cells flagged unreliable.

Separate from estimate() on purpose. The data layer returns everything with a reason attached; discarding is the caller’s decision, and a silent drop at fetch time would hide how much of a cross-tabulation is too thin to use.