Distributions#

T.reference is the reference distribution. To obtain a distribution in source space, supply a map to T.predictive_distribution. The guide walks through a complete example.

Construct a distribution#

yemale.ot.reference(n, dimension)[source]#

Return a shared Reference with n + 1 cells in dimension dimensions.

n is the number of calibration observations, not the number of cells.

Transport.reference_distribution(point)[source]#

Return the reference distribution within one candidate’s assigned cell.

\[K(z,\cdot)=\nu(\,\cdot\mid L_{k(z)}).\]

Samples are reference points, not predictions in source coordinates. Accept one candidate; return a Law. Requires a Reference target.

Transport.predictive_distribution(map_from_reference, *, inverse=None, inverse_logabsdet=None)[source]#

Create a predictive Law by choosing how to fill each source cell.

Write Q_j for map_from_reference on reference cell L_j. Map each L_j into its matching source cell; each gets probability 1 / (n + 1).

\[\Pi^Z=\frac1{n+1}\sum_{j=1}^{n+1} (Q_j)_\#\nu(\,\cdot\mid L_j).\]

This is the law of Q_j(U) after drawing a cell uniformly and U within it. Q_j chooses a distribution inside the cell; it is not an inverse of T.

Parameters:
  • map_from_reference – Called as map_from_reference(points, labels) with reference points (q, d) and cell labels (q,). Returns points (q, d) in the matching source cells. For density, use a differentiable, one-to-one map with nonsingular Jacobian.

  • inverse – Optional inverse called as inverse(source_points, cell_labels). Supply with inverse_logabsdet for density and entropy. Returns (q, d).

  • inverse_logabsdet – Optional callable with the same inputs as inverse. Returns the inverse Jacobian’s log absolute determinant, as a scalar, (q,), or (q, 1).

Notes

Callbacks must treat input points as read-only. Outside the mapped support, the inverse may return finite placeholders paired with a -inf log determinant.

Reference distribution#

n is the number of fitted source points and dimension is the number of coordinates. centers, ranks, and signs have one entry per reference cell.

class yemale.ot.Reference(n, dimension)[source]#

A distribution on the unit ball, split into n + 1 equal-probability cells.

The direction is uniform on the sphere, independently of a radius uniform on [0, 1]. In one dimension, this is uniform on [-1, 1]. n is the calibration size; the extra cell accounts for the candidate. Write nu for this law, L_j for its cells, and d for dimension:

\[U=R\Theta,\quad R\sim\mathrm{Uniform}[0,1],\quad \Theta\sim\mathrm{Uniform}(\mathbb S^{d-1}),\quad R\perp\Theta, \qquad \nu(L_j)=\frac1{n+1}.\]
n#

Source size; the reference contains n + 1 cells.

Type:

int

dimension#

Number of coordinates.

Type:

int

property centers#

Mean reference point in each cell, with shape (n + 1, dimension).

\[m_j=\mathbb E_\nu[U\mid U\in L_j].\]
property ranks#

Distance of each cell centre from the origin, with shape (n + 1,).

property signs#

Unit directions of centers, with the same shape; zero stays zero.

sample(size=1, *, rng=None)[source]#

Draw (size, dimension) samples; radius is uniform, not volume.

rng is a seed or NumPy generator; omit it for fresh randomness.

locate(value)[source]#

Return cell labels, choosing the smallest on shared boundaries.

Points outside the unit ball return -1.

logpdf(value)[source]#

Return the natural logarithm of pdf(value).

Accept (d,) or (..., d); return one value per point. Return -inf outside the ball and +inf at the origin for d > 1.

pdf(value)[source]#

Return density per unit volume at (d,) or (..., d) points.

\[p_\nu(u)=\frac{\Gamma(d/2)}{2\pi^{d/2}}\|u\|^{1-d}, \qquad 0<\|u\|\leq1.\]

Return one value per point, zero outside the unit ball. At the origin the density is infinite for d > 1; in d = 1 it is 1/2.

density_region(mass, *, n_integration_points=256)[source]#

Select a reference-density superlevel set by approximate probability mass.

Include all ties. n_integration_points sets the points per cell. Return a DensityRegion with contains, threshold, and mass.

expect(function, *, n_integration_points=64, rng=None)[source]#

Approximate E[function(U)] for a reference point U.

\[\mathbb E_\nu[f(U)]=\int f(u)\,\nu(du).\]

function receives points (q, d) and returns values (q, ...). The result averages over points and keeps the remaining axes. n_integration_points is the number of points per cell. rng=None uses fixed points (interval midpoints in 1-D, Halton points otherwise). A seed or NumPy generator draws random points within each cell.

moment(powers)[source]#

Return an exact raw moment, with one power per coordinate.

powers gives the nonnegative integers in alpha. For example, moment((2, 0)) means E[U[0]**2]. Any odd power gives zero. For even powers, with \(|\alpha|\) their sum,

\[\mathbb E_\nu\!\left[\prod_{r=1}^d U_r^{\alpha_r}\right] =\frac{1}{|\alpha|+1} \frac{\Gamma(d/2)}{\Gamma((d+|\alpha|)/2)} \prod_{r=1}^d\frac{\Gamma((\alpha_r+1)/2)}{\Gamma(1/2)}.\]
mean()[source]#

Return the mean reference point: a zero vector of length dimension.

\[\mathbb E_\nu[U]=0.\]
covariance()[source]#

Return the covariance: identity divided by 3 * dimension.

\[\mathrm{Cov}_\nu(U)=\mathbb E_\nu[UU^\top]=\frac{I_d}{3d}.\]

The result has shape (dimension, dimension).

entropy()[source]#

Return differential entropy -E[log(pdf(U))], in nats.

\[h(\nu)=\log\left(\frac{2\pi^{d/2}}{\Gamma(d/2)}\right)-(d-1).\]

Cell mixtures and mapped distributions#

reference is the underlying Reference. cells are its selected zero-based labels and weights are their probabilities.

class yemale.ot.Law(reference, cells=None, weights=None, forward=None, backward=None)[source]#

A probability distribution built from cells of a Reference.

Sampling chooses a cell using its weight, then draws from the reference distribution within that cell. An optional map transforms the samples.

reference#

Reference distribution containing the selected cells.

Type:

yemale.ot.reference.Reference

cells#

Distinct zero-based reference-cell labels.

Type:

numpy.ndarray

weights#

Cell probabilities, aligned with cells and summing to one.

Type:

numpy.ndarray

Choose reference cells, their probabilities, and optional maps.

Parameters:
  • reference – Reference containing the cells.

  • cells – Distinct integer labels in [0, reference.n]. Defaults to all.

  • weights – Nonnegative probabilities, one per selected cell, summing to one. Defaults to equal weights over the selected cells.

  • forward – Optional forward(points, labels) taking (q, d) reference points and (q,) labels. Returns mapped points (q, d). For density evaluation, it must be one-to-one across selected cells.

  • backward – Optional backward(points) taking mapped points (q, d). Returns reference points (q, d) and the inverse Jacobian’s log absolute determinant, as a scalar, (q,), or (q, 1). Required for density and entropy when forward is set.

Notes

Callbacks must treat input points as read-only. A -inf log determinant can mark points outside the mapped support.

property dimension#

Number of coordinates in each sample.

sample(size=1, *, rng=None)[source]#

Draw (size, dimension) samples; rng is a seed or NumPy generator.

logpdf(value)[source]#

Return the natural logarithm of pdf(value).

Accept a point (d,) or batch (..., d); return one value per point. A mapped law requires backward; outside its support, return -inf.

pdf(value)[source]#

Return density per unit volume, not the probability of a cell.

Accept (d,) or (..., d); return one value per point, zero outside the support. For reference law nu, n = reference.n, cell weight w_j, and inverse map q_j,

\[p(x)=(n+1)w_j p_\nu(q_j(x))|\det Dq_j(x)|.\]

The formula holds almost everywhere on the image of cell j. Mapped laws require backward and a differentiable map with nonsingular Jacobian, one-to-one across selected cells. Without a map, q_j is the identity.

density_region(mass, *, n_integration_points=256)[source]#

Select a density superlevel set with approximately the requested mass.

\[A_t=\{x:p(x)\ge t\},\qquad \int_{A_t}p(x)\,dx\approx\mathrm{mass}.\]

mass is in (0, 1]. Include all ties at the cutoff; the returned mass may therefore be larger. This is not a future-data coverage guarantee. n_integration_points sets the points per selected cell; the prepared density values are reused for subsequent mass requests. Mapped laws require the inverse used by logpdf.

expect(function, *, n_integration_points=64, rng=None)[source]#

Approximate E[function(X)] under this distribution.

\[\mathbb E[f(X)] =\sum_j w_j\,\mathbb E_\nu[f(Q_j(U))\mid U\in L_j].\]

Here nu is the reference law, w_j the cell weight, and Q_j the forward map in cell L_j (the identity when no map is supplied).

function receives points (q, d) and returns values (q, ...). The result averages over points and keeps the remaining axes. n_integration_points is the number of points per selected cell. rng=None uses fixed points (interval midpoints in 1-D, Halton points otherwise). A seed or NumPy generator draws random points within each cell.

moment(powers, *, n_integration_points=64)[source]#

Return a raw moment, with one power per coordinate.

powers gives the nonnegative integers in alpha. For example, moment((2, 0)) means E[X[0]**2].

\[\mathrm{moment}(\alpha) =\mathbb E\!\left[\prod_{r=1}^d X_r^{\alpha_r}\right].\]

Reference-cell moments are exact. Mapped moments use fixed integration points; n_integration_points sets their number per cell.

mean(*, n_integration_points=64)[source]#

Return the mean E[X] as a vector of length dimension.

\[\mathbb E[X]=(\mathbb E[X_1],\ldots,\mathbb E[X_d])^\top.\]

Reference-cell means are exact. Mapped means use fixed points; n_integration_points sets their number per cell.

covariance(*, n_integration_points=64)[source]#

Return the covariance matrix, with shape (dimension, dimension).

\[\mathrm{Cov}(X)=\mathbb E[(X-\mathbb E[X])(X-\mathbb E[X])^\top].\]

Reference-cell covariances are exact. Mapped covariances use fixed integration points; n_integration_points sets their number per cell.

entropy(*, n_integration_points=64)[source]#

Return differential entropy -E[log(pdf(X))], in nats.

\[h(X)=-\int p(x)\log p(x)\,dx.\]

Reference-cell mixtures use an exact formula, including the cell weights. Mapped laws require backward and integrate -logpdf using n_integration_points fixed points per cell.