Conformal prediction#

Build prediction regions with scalar or vector-valued scores, and choose a predictive law for sampling and summaries. Start with a fitted model or with prediction arrays and their observed outcomes.

Use your model#

Given a fitted model and calibration data not used to train it:

from yemale import extend

model = extend(model)
model.conformalize(X_cal, y_cal)
cpd = model.predict_distribution(X_new)

region = cpd.region(0.9)
region.contains(y_new)  # Test outcomes, paired with the new inputs.
region.coverage        # 0.9

model.predict keeps returning point predictions. After changing or retraining the model, calibrate again. Regions are randomized by default; pass rng=0 to region for reproducibility. Membership stays fixed once the region is built.

Use prediction arrays#

The model wrapper calls the same array-based API. For example, with scalar predictions and observed outcomes:

import numpy as np
from yemale import conformalize

rng = np.random.default_rng(0)
predictions = rng.normal(size=99)
outcomes = predictions + rng.normal(scale=0.5, size=99)

cp = conformalize(predictions, outcomes)
cpd = cp.predict([0.2, 1.0])
region = cpd.region(0.9, rng=0)
region.contains([0.3, 1.2])  # One outcome for each prediction.

candidates = np.linspace(-2, 2, 100)
region[0].select(candidates)  # Keep candidates for the first prediction.

Calibration accepts (n,) scalar arrays or (n, d) vector arrays. A prediction batch has one region per row; indexing selects a region, and select returns the original candidate outcomes.

Scores#

The default score is outcomes - predictions. Supply score= to conformalize or extend for another batch function: it receives predictions and outcomes and returns one numeric scalar or vector score per pair. The transport is fitted once to those scores and reused for every prediction.

Sampling and summaries#

A score law describes the distribution of scores; the inverse score converts its samples into outcomes. For residuals, this simply adds the prediction. For one prediction, supply candidate= to build a smooth law from the calibration scores and that candidate’s score:

cpd = cp.predict(0.2, candidate=0.3)
cpd.sample(1000, rng=0)
cpd.mean()
cpd.cov()

The candidate stays fixed while samples vary. This path uses the default OT temperature and provides sampling, expectations and moments. For a custom score, supply its forward and inverse functions in a ScoreMap.

To choose the score distribution yourself, pass law= through cp.predict(prediction, law=score_law) or model.predict_distribution(inputs, law=score_law). Density and entropy use this path with the density callbacks.

Coverage#

Keep the predictor and score fixed independently of calibration. Under the exchangeability and augmented-assignment uniqueness assumptions in the guide, coverage averages over calibration samples, future observations and the region’s randomization. This is marginal coverage.

Shapes#

At prediction time, a vector (width,) binds one prediction; a matrix (batch, width) keeps a prediction-batch axis. Scalars accept 3 or [3] for one prediction and [3, 5] for two. [[3]] keeps a one-row batch axis.

Membership pairs prediction and outcome rows. Use cpd[i] to select one prediction and test a collection of candidates. Geometric readouts keep an outcome row axis: for q outcomes and s score coordinates, transform and sign return (q, s); rank, potential and region.contains return (q,).

For a custom shape, cpd.region(reference_set=B) uses the same reference-space predicate as cp.transport.region(B). See OT regions.

Predictive laws#

In the OT layer, the fitted source points are the calibration scores. A score law \(\Pi\) becomes an outcome law through the inverse score:

\[\Pi_x=(S_x^{-1})_\#\Pi, \qquad \mathbb E_{\Pi_x}[f(Y)] =\mathbb E_\Pi[f(S_x^{-1}(Z))].\]

The candidate=y shortcut uses the OT smooth pullback at \(S_x(y)\):

\[\Pi_{x,\tau}(\,\cdot\,;y) =\left[S_x^{-1}\circ Q_\tau(\,\cdot\,;S_x(y))\right]_\#\nu.\]

Here \(y\) is fixed across draws. To specify a law within each hard score cell, use cp.transport.predictive_distribution and pass the result as law=. The uniform-gap example shows a complete construction. Equal-cell calibration requires mass \(1/(n+1)\) in every fitted score cell; the smooth construction need not preserve these masses. Regions are determined directly by the fitted transport, independently of the sampling law.

Supply ScoreMap.inverse on the chosen law’s support, preserving coordinate dimension. Batched summaries keep a leading prediction axis; cpd.sample(size) returns (batch, size, d) for outcome width d.

Density and entropy#

The forward density uses the smooth transport and the score Jacobian:

\[p_{\tau,x}(y)=p_\nu(T_\tau(S_x(y))) \left|\det DT_\tau(S_x(y))\right| \left|\det D_yS_x(y)\right|.\]
density = cpd.density(outcomes)

This needs no candidate or inverse score. The residual score already provides the Jacobian; custom scores provide ScoreMap.jacobian and must be one-to-one on the evaluation domain. log_density returns the same readout in log scale. Sampling and moments still use \(Q_\tau\) through Law. The forward density is not generally normalized or exactly the PDF of those samples.

For pdf, logpdf and entropy, supply a score law with density callbacks through law=, and the score Jacobian through ScoreMap.jacobian. The residual score already provides this Jacobian. For mutually inverse differentiable score maps with nonsingular outcome Jacobian,

\[p_{\Pi_x}(y)=p_\Pi(S_x(y))\left|\det D_y S_x(y)\right|.\]

Evaluate densities on a domain where these inverse conditions hold. For a mapped score law, its density also requires the inverse of its reference-to-score map and that inverse’s log-Jacobian, as in the guide. Entropy has the same requirements.

Use one prediction, or cpd[i], for pdf, logpdf and density_region(mass=0.9). The latter selects probability mass under the supplied law; cpd.region(0.9) targets marginal coverage of future outcomes. The forward readouts density and log_density also accept paired batches.

Potential and gradient#

cpd.potential(outcomes) evaluates the fitted transport potential through the score. cpd.potential_gradient(outcomes) uses ScoreMap.jacobian to express its gradient in outcome coordinates. Both operate directly on the transport.

API#

yemale.conformalize(predictions, outcomes, *, score=residual, target=None)[source]#

Fit once to scores of paired calibration predictions and outcomes.

Parameters:
  • predictions – Array (n, p); scalar predictions may use (n,).

  • outcomes – Array (n, d); scalar outcomes may use (n,).

  • score – ScoreMap or batch callable. Defaults to outcomes - predictions. Return one score per pair, with shape (n, s) or (n,) for scalar scores. Only these scores must be numeric; s may differ from the prediction and outcome widths.

  • target – Optional target passed to yemale.ot.fit().

Return a Conformalizer. The predictor and score must be fixed independently of calibration. DH coverage assumptions apply to the resulting scores.

yemale.extend(model, *, score=residual)[source]#

Add conformal readouts to a fitted predictor without changing its methods.

Fit the model before extending it. Delegated methods keep their return values: a model’s self-returning fit returns that model, not this wrapper. score(predictions, outcomes) is evaluated on whole arrays.

class yemale.conformal.ScoreMap(forward, inverse=None, jacobian=None)[source]#

A score S_x and its optional inverse and outcome derivative.

Callbacks receive the fixed model prediction at x. Write p, d, and s for the prediction, outcome, and score widths, and q for the number of outcomes.

Parameters:
  • forward – forward(predictions, outcomes) returns scores (q, s) or (q,) for scalar scores. Outcomes have shape (q, d); predictions have shape (q, p) or (1, p) when one prediction accompanies many outcomes.

  • inverse – Optional inverse(predictions, scores) returning outcomes (q, d). Used to convert score samples into outcome samples. It must invert the score on the chosen law’s support.

  • jacobian – Optional jacobian(predictions, outcomes) returning D_y S_x with shape (s, d) or (q, s, d). Required for potential_gradient and density. Density requires s = d and mutually inverse, differentiable score maps with nonsingular Jacobian at every queried outcome. For a restricted inverse branch, evaluate density on that branch’s domain.

class yemale.conformal.Conformalizer[source]#

A fitted score transport; use cp.predict(prediction) for its readouts.

Create with yemale.conformalize(predictions, outcomes).

transport#

Shared DH Transport fitted to the calibration scores.

Type:

yemale.ot.transport.Transport

score#

ScoreMap used for calibration and subsequent outcomes.

Type:

yemale.conformal.predict.ScoreMap

predict(prediction, *, candidate=None, law=None)[source]#

Bind one prediction or a batch, without refitting; return a CPD.

For prediction width p, (p,) is one prediction and (batch, p) is a batch. For p = 1, a scalar or [value] is one prediction; a longer vector is a batch of scalar predictions. [[value]] is a one-member batch, so select cpd[0] for single-prediction density.

Regions, ranks and the forward density readout use the fitted transport directly. For sampling and summaries, choose one of:

  • candidate: one outcome fixing the augmented smooth law at S_x(candidate), using the default temperature. Requires reference cells and an inverse score. Supports sample, expect, mean, cov and moment; pdf, logpdf and entropy require the law path below.

  • law: a score-space Law Pi of your choice. The inverse score converts its samples to outcomes. pdf and logpdf also require the law’s density callbacks and the score Jacobian.

The candidate stays fixed across draws. For equal-cell calibration, choose a law assigning mass 1 / (n + 1) to each fitted score cell.

class yemale.conformal.CPD[source]#

Conformal predictive distribution (CPD) and readouts at given predictions.

Create with cp.predict(prediction). Queries pair prediction and outcome rows, or use one prediction for many outcomes. cpd[i] selects a batch member and shares the fitted transport. Geometric queries keep the outcome row axis: for q outcomes, s score coordinates, and d outcome coordinates, transform and sign return (q, s), rank and potential return (q,), and potential_gradient returns (q, d). A single outcome has q = 1.

Sampling and summaries need an inverse score and either a candidate outcome or an explicit score law. Their shapes are documented on each method.

property transport#

The shared DH transport on calibration scores.

__getitem__(index)[source]#

Select predictions from a batch, sharing its transport and score law.

transform(outcomes)[source]#

Map outcomes to reference coordinates, one vector per outcome.

\[G_x(y)=T(S_x(y)).\]

The vector’s norm is rank; its unit direction is sign.

rank(outcomes)[source]#

Return one center-outward rank per outcome, on the reference scale.

\[\operatorname{Rank}_x(y)=\operatorname{Rank}(S_x(y)).\]
sign(outcomes)[source]#

Return each outcome’s reference-space direction; zero maps give zero.

\[\operatorname{Sign}_x(y)=\operatorname{Sign}(S_x(y)).\]
potential(outcomes)[source]#

Evaluate the potential through the score, one value per outcome.

\[y\mapsto\Phi(S_x(y)).\]
potential_gradient(outcomes)[source]#

Return the potential gradient in outcome coordinates.

\[H_x(y)=DS_x(y)^\top T(S_x(y)).\]

This equals the gradient of potential where the score and DH potential are differentiable. At ties, it uses the DH-selected target. Requires ScoreMap.jacobian; returns one vector per outcome.

region(coverage=None, *, reference_set=None, randomized=True, rng=None)[source]#

Return an outcome region by composing a DH region with S_x.

For a hard reference set B:

\[\mathcal R_B(x)=\{y:T(S_x(y))\in B\}.\]

Supply coverage (default 0.9) or reference_set, not both. Reference sets are predicates on target coordinates, as in T.region. With coverage, randomization is on by default: draw once inside each reference cell and include its source cell if the draw lies inside the reference ball. This selection stays fixed across membership calls. Use rng to reproduce it, or randomized=False to select by cell centres instead, with whole-shell rounding as in T.quantile_region.

reference_distribution(outcome)[source]#

Return K_x(y, .) = K(S_x(y), .) for one candidate outcome.

The returned Law draws reference points, not future outcomes.

log_density(outcomes, *, temperature=None)[source]#

Return the forward smooth log-density readout, one value per outcome.

Requires ScoreMap.jacobian and equal score/outcome dimensions. No candidate or inverse score is needed. temperature is passed to transport.smooth; prediction batches use paired outcomes.

density(outcomes, *, temperature=None)[source]#

Evaluate the forward smooth density through the score.

\[p_{\tau,x}(y)=p_\tau(S_x(y))|\det D_yS_x(y)|.\]

Here p_tau is transport.smooth(temperature).density. Evaluate on a domain where the score is one-to-one and differentiable. Return one value per outcome, as in log_density. This readout need not integrate to one or equal the PDF of the chosen sampling law.

sample(size=1, *, rng=None)[source]#

Draw outcomes: (size, d) or (batch, size, d) for a CPD batch.

expect(function, *, n_integration_points=64, rng=None)[source]#

Compute E_Pi_x[f(Y)] = E_Pi[f(S_x^{-1}(Z))] through the chosen Law.

function receives (q, d) outcomes. Numerical integration and its controls are those of yemale.ot.Law.expect().

mean(*, n_integration_points=64)[source]#

Return the outcome mean, shape (d,) or (batch, d).

cov(*, n_integration_points=64)[source]#

Return outcome covariance, shape (d, d) or (batch, d, d).

moment(powers, *, n_integration_points=64)[source]#

Return E[prod(Y**powers)] using the chosen Law’s moment method.

entropy(*, n_integration_points=64)[source]#

Return outcome differential entropy in nats; requires density.

logpdf(outcomes)[source]#

Return outcome log density for one prediction; index a CPD batch first.

A point (d,) returns a scalar; a batch (*batch_shape, d) returns batch_shape. For d = 1, a scalar or one-element vector is a point, and a longer vector is a batch. Density does not keep a singleton outcome axis unless it is present in the input batch shape.

pdf(outcomes)[source]#

Return outcome density for one prediction; requires the score Jacobian.

Input and output shapes follow logpdf; index a CPD batch first.

density_region(mass, *, n_integration_points=256)[source]#

Return a density-mass region of one completed outcome law, not coverage.

class yemale.conformal.Region[source]#

Candidate outcomes whose scores belong to a shared DH region.

Create with cpd.region(coverage) or model.predict_region(inputs, coverage).

property coverage#

Marginal coverage reported by the underlying DH region.

contains(outcomes)[source]#

Return one Boolean per outcome, pairing rows for a prediction batch.

select(candidates)[source]#

Filter candidates for one region, preserving their values and order.

For a batch, select one region with regions[i] first.

__getitem__(index)[source]#

Select predictions, retaining the same DH region and randomization.

class yemale.conformal.Extended[source]#

Add calibration and regions to a predictor; delegate its other methods.

Create with yemale.extend(model) after fitting the predictor.

conformalize(inputs, outcomes, *, target=None, **predict_kwargs)[source]#

Calibrate from predictions on observations not used to train the model.

predict_distribution(inputs, *, candidate=None, law=None, **predict_kwargs)[source]#

Bind the calibrated readouts to a batch of new predictions.

For one prediction, supply a candidate outcome for smooth sampling and moments. Alternatively, supply a score-space law. Without either, the CPD supports regions and geometric readouts.

predict_region(inputs, coverage=None, *, reference_set=None, randomized=True, rng=None, **predict_kwargs)[source]#

Return regions for new inputs, by coverage or a reference-set predicate.