pyforestry.base package#

Subpackages#

Submodules#

pyforestry.base.aggregation module#

The single plot-to-stand estimator.

Every per-hectare stand figure in this package comes from here: Stand reports its metrics from aggregate_plots(), and so does the working copy a SimulationContext holds. They used to carry an implementation each, and the two disagreed – the context averaged a species over only the plots where it occurred, so a two-plot stand with one spruce plot and one pine plot read 1.0 stems/ha through Stand and 2.0 through the context. One estimator makes that class of divergence unrepresentable.

The estimator’s conventions, in one place:

  • A plot’s trees are expanded over its effective area, area_ha * (1 - occlusion), so a partly occluded plot represents the area actually searched.

  • Tree.weight_n is honoured everywhere: a record standing for ten stems counts as ten.

  • A species absent from a plot contributes an explicit zero for that plot, so the cross-plot mean divides by the number of plots – not by the number of plots where the species happens to occur.

  • A record with no diameter is still a stem, but carries no basal area and is excluded from the diameter-derived metrics (QMD, BAWAD) so their numerator and denominator describe the same trees.

  • Every precision is the standard error of the stand mean over plots. It is a between-plot error only and includes no measurement or model error, so a metric built on curve-imputed heights still reports a measured-quality error.

  • Ratio metrics (Lorey’s mean height, the basal-area weighted diameter, QMD) are reduced from the per-plot series through ratio_and_se(), which carries the correlation between numerator and denominator rather than assuming it away.

pyforestry.base.aggregation.COMPONENT_KEYS = ('stems', 'stems_d', 'ba', 'd3', 'gh', 'ba_h')#

Per-plot, per-species quantities accumulated in one pass and reused to derive every metric: stem count, stem count of trees carrying a diameter, basal area, sum of cubed diameters (the BAWAD numerator), and the basal-area weighted height sum with its matching basal area (Lorey’s).

class pyforestry.base.aggregation.PlotAggregation(stems: TreeName | str, ~pyforestry.base.helpers.primitives.area_aggregates.Stems]=<factory>, basal_area: TreeName | str, ~pyforestry.base.helpers.primitives.area_aggregates.StandBasalArea]=<factory>, bawad: TreeName | str, ~pyforestry.base.helpers.primitives.bawad.BasalAreaWeightedDiameter]=<factory>, loreys_height: TreeName | str, ~pyforestry.base.helpers.primitives.loreys_mean_height.LoreysMeanHeight]=<factory>, species_components: Any, ~typing.Dict[str, ~typing.List[float]]]=<factory>, total_components: Dict[str, ~typing.List[float]]=<factory>, species_order: TreeName, ...]=(), missing_diameter: int = 0)[source]#

Bases: object

Per-hectare stand metrics estimated as the mean over a set of plots.

Variables:
missing_diameter: int = 0#
property qmd: QuadraticMeanDiameter#

Stand quadratic mean diameter, reduced from the total component series.

species_order: Tuple[TreeName, ...] = ()#
pyforestry.base.aggregation.aggregate_plots(plots: Iterable['CircularPlot'], *, warn_missing_diameter: bool = True) → PlotAggregation[source]#

Estimate per-hectare stand metrics as the mean over plots.

One pass over every plot’s trees accumulates, per species, the per-hectare stem count, basal area, sum of cubed diameters and basal-area weighted height sum. See the module docstring for the conventions this applies.

Parameters:
  • plots – The plots to aggregate. They form one population; a caller with several independent populations aggregates each separately.

  • warn_missing_diameter – Whether to warn about tree records carrying no diameter. A simulation refreshing its metrics every step passes False: the stand it was built from already reported this once, and repeating it per step is noise rather than information. The count is reported on the result either way.

Returns:

A PlotAggregation holding the metrics and the per-plot series they were reduced from.

pyforestry.base.aggregation.mean_and_se(values: Sequence[float]) → Tuple[float, float][source]#

Return the sample mean and standard error of the mean of values.

The standard error is the sample standard deviation divided by sqrt(n); it is 0.0 when fewer than two values are supplied.

pyforestry.base.aggregation.metrics_from_series(series: Dict[str, List[float]]) → Tuple[Dict[str, float], Dict[str, float]][source]#

Derive the stand metrics and their standard errors from component series.

series maps each per-plot component accumulated by aggregate_plots() (stems, ba, d3, gh, ba_h) to its per-plot values. Additive metrics come from the plot-to-plot mean; the ratio metrics (basal-area weighted diameter, Lorey’s mean height) go through ratio_and_se().

Returns a (value, precision) pair of dicts keyed by metric name.

pyforestry.base.aggregation.qmd_from_components(ba_mean: float, ba_se: float, stems_mean: float, stems_se: float) → QuadraticMeanDiameter[source]#

Quadratic mean diameter (cm) from basal-area and stem means, with its SE.

QMD = sqrt(40000 * BA / (pi * N)); the standard error is propagated from the basal-area and stem standard errors by the delta method.

This form treats the two as independent, which overstates the error because basal area and stem count rise and fall together across plots. It is only for callers that have nothing but the reduced means – angle-count stands, where basal area comes from the tally and stems from the recorded diameters. Where the per-plot series survive, qmd_from_series() is both covariance-aware and numerically better behaved.

pyforestry.base.aggregation.qmd_from_series(series: Dict[str, List[float]]) → QuadraticMeanDiameter[source]#

Quadratic mean diameter from per-plot basal-area and stem series.

Uses stems_d – the stems that actually carry a diameter – so that the numerator and denominator describe the same trees; a record with no diameter contributes no basal area and must not enter the count either.

QMD = sqrt(40000/pi * R) for the mean basal area per stem R, so the standard error is obtained by propagating through that square root from ratio_and_se(). Going via the ratio rather than combining the two standard errors and their covariance term by term avoids the cancellation those three near-equal terms suffer: a stand whose plots share one QMD reports exactly zero here instead of a small numerical residue.

pyforestry.base.aggregation.ratio_and_se(numerator: Sequence[float], denominator: Sequence[float]) → Tuple[float, float][source]#

Ratio of two per-plot means with a covariance-aware standard error.

Both arguments are the per-plot series of the numerator and denominator (not their means), because a ratio metric such as Lorey’s mean height or the basal-area weighted diameter draws both parts from the same trees: they are strongly correlated, and treating them as independent invents variance that is not there.

The standard error is the classical ratio-estimator form, which carries that correlation exactly:

SE(R) = sqrt(Var(N_i - R * D_i) / n) / mean(D)

A stand whose plots all share the same ratio therefore reports a standard error of zero, however much the plots differ in density.

Returns (0.0, 0.0) when the denominator mean is non-positive (e.g. no basal area, or no measured heights for Lorey’s mean).

pyforestry.base.aggregation.sum_series(components: Dict[Any, Dict[str, List[float]]], species: Iterable[Any]) → Dict[str, List[float]][source]#

Elementwise per-plot sum of the component series over species.

Summing the series before reducing keeps a species group covariance-aware in the same way the stand total is: species that trade off against each other across plots do not each contribute independent variance.

pyforestry.base.contracts module#

Foundational provenance and introspection contracts.

The SourceReference dataclass and the Describable / FormulaModuleDescriptor protocols describe what a component is — its identity, bibliographic provenance, species applicability, and units — independent of any simulation runtime. They live at the base level because formula/domain modules across every region expose them (for example a module’s DESCRIPTOR), so nothing above base should have to reach into the simulation package for this vocabulary.

This is the only import path for these names. pyforestry.simulation.contracts holds what the runtime defines (SimulationPreset) and deliberately does not re-export this vocabulary, so that an equation module never has to import from the simulation tier to describe itself.

class pyforestry.base.contracts.Describable(*args, **kwargs)[source]#

Bases: Protocol

Any component that declares its identity and provenance.

property component_id: str#

Stable identifier for this component.

property source: SourceReference#

Bibliographic provenance for this component.

class pyforestry.base.contracts.FormulaDescriptor(component_id: str, source: SourceReference, species_groups: Mapping[str, frozenset[str]]=<factory>, units: Mapping[str, str]=<factory>, kernel_names: Tuple[str, ...]=(), kind: str = 'formula', domain: str | None = None, composes: Tuple[str, ...]=())[source]#

Bases: object

Concrete, declarative FormulaModuleDescriptor for formula modules.

A module publishes its introspection metadata with minimal boilerplate:

DESCRIPTOR = FormulaDescriptor(
    component_id="brandel_1990_volume",
    source=SourceReference(author="Brandel, G.", year=1990, title="..."),
    kernel_names=("BrandelVolume",),
)

species_groups should use full scientific names (e.g. "Pinus sylvestris") so the catalog’s species filter matches cleanly.

kind distinguishes the two tiers the catalog surfaces: "formula" for equation/kernel modules (the default) and "model" for composed, runnable models (the adapters/ and systems/ tiers). A "model" descriptor may set domain to the scientific domain it belongs to (e.g. "growth", since its module path lives under adapters/) and list the formula component_id``s it ``composes.

composes: Tuple[str, ...] = ()#
domain: str | None = None#
kernel_names: Tuple[str, ...] = ()#
kind: str = 'formula'#
class pyforestry.base.contracts.FormulaModuleDescriptor(*args, **kwargs)[source]#

Bases: Protocol

Introspection contract for a formula module (not per-function).

Each formula package exposes a module-level DESCRIPTOR object implementing this protocol – in practice a FormulaDescriptor. The descriptor sits next to the kernel functions, not around them; kernel function signatures are not changed.

kind, domain and composes are part of the contract because pyforestry.catalog reads them. They were absent from this Protocol while the catalog read all three with getattr defaults, so a descriptor could satisfy the declared contract in full and still be unable to say it was a model rather than a formula.

property component_id: str#

Stable identifier for this formula module.

property composes: Sequence[str]#

component_id``s of the formula modules a ``"model" descriptor composes.

property domain: str | None#

Scientific domain, when the module path does not already say it.

property kernel_names: Sequence[str]#

Public function names exposed by this formula module.

property kind: str#

"formula" for an equation module, "model" for a composed model.

property source: SourceReference#

Bibliographic provenance for this formula module.

property species_groups: Mapping[str, frozenset[str]]#

Species groups and their constituent species identifiers.

property units: Mapping[str, str]#

Unit contract mapping parameter/output names to unit strings.

class pyforestry.base.contracts.ImputedValue(attribute: str, value: float, imputer_id: str, source: SourceReference)[source]#

Bases: object

A modelled value for one attribute, with its provenance.

Lives here rather than next to the imputers because Tree stores these, and a leaf data class must not depend on the imputation package that fills them in.

Variables:
  • attribute (str) – The attribute this value stands in for, e.g. "height_m".

  • value (float) – The modelled value, in the attribute’s own units.

  • imputer_id (str) – component_id of the imputer that produced it.

  • source (pyforestry.base.contracts.SourceReference) – The publication behind that imputer. A user-supplied callable carries the "(none)"/year-0 sentinel, so an uncited value is visibly uncited rather than silently unattributed.

property is_cited: bool#

Whether this value traces to a publication rather than a bare callable.

class pyforestry.base.contracts.SourceReference(author: str, year: int, title: str, appendix: str = '', note: str = '')[source]#

Bases: object

Bibliographic provenance for a scientific formula or model.

appendix: str = ''#
note: str = ''#

Module contents#