Glossary

aleatoric uncertainty

In mc_fit, aleatoric uncertainty denotes the uncertainty arising from non-uniqueness of the inversion problem. It is quantified from the scatter of central models whose misfit lies within the imprecision interval of the best central model. This usage is broader than the conventional statistical meaning of aleatoric uncertainty, which is usually restricted to uncertainty arising solely from inherently random processes.

best central model

The central model that, among all successful tries made during the central model analysis, has either the minimum misfit (default) or maximum Bayes score. Whether misfit or Bayes score is used as the selection criterion is specified by the Bayes option.

best overall model

The model having the lowest objective function value found during the perturbation analysis, provided it has a lower objective function value than the best central model. The best overall model is identified separately only to indicate that the perturbation analysis found a better-fitting model than the best central model, thereby illustrating the range of objective function values generated by the perturbation analysis.

central model

A forward model generated as part of the central model analysis, in which all known parameters are fixed at their central (most probable) values. In mc_fit, the known parameters typically comprise the analytical and thermodynamic data. The population of central models samples the misfit function over the specified inversion parameter ranges. After the misfit imprecision has been estimated by the perturbation analysis, the retained central models characterize the aleatoric uncertainty of the inverse solution.

central model analysis

The first stage of an mc_fit inversion. During this stage, the optimization is repeated from randomly generated starting values of the inversion parameters using the central (most probable) values of the analytical and thermodynamic data. The analysis identifies the best central model and provides the population of central models. The retained central models, identified by applying the misfit imprecision filter determined by the perturbation analysis, characterize the aleatoric uncertainty of the inverse solution.

composant

A phase with the same composition as a saturated component. If multiple component saturation constraints have been imposed, then the composant of the n’th saturated component is a phase that must contain component n, and may contain components 1 to n - 1, but does not contain any additional components ([Connolly1990]).

compositional variable

An intensive state variable (e.g., molar fractions, molar properties, densities, specific properties, chemical compositions) that is a ratio of extensive variables ([Hillert1986], [Connolly1990]).

data uncertainty

The component of total uncertainty in the inversion parameters arising from uncertainty in the analytical and thermodynamic data. In mc_fit, it is estimated from the covariance of inversion results obtained during the perturbation analysis.

ds5/ds6

THERMOCALC (aka TC or HPx-eos) thermodynamic dataset major version designations. Until 2016, Perple_X implementations of the datasets named the thermodynamic files by date (e.g., hp98ver.dat) rather than TC version. In Perple_X documentation, all TC datasets prior to hp11ver.dat are loosely referred to as ds5 datasets ([Holland1998]). TC datasets from hp11ver.dat onward are ds6 datasets ([Holland2011]). From 2015 onward, Perple_X implementations of TC datasets have been named using the TC version number (e.g., hp622ver.dat corresponds to TC version 6.22).

excess oxygen

The amount of oxygen component that must be added or subtracted from an oxide or metallic component to define the actual redox state of the metal represented by the oxide or metallic component (Appendix D).

fit-all-data criterion

A criterion that requires a model to fit all analytical data within a specified uncertainty interval. The criterion may be applied in either mc_fit or mc_fit_plot to filter the population of central models.

In mc_fit:

  • When must_fit_all_data is T only models that meet the criterion are output to the my_project_central.pts file. When must_fit_all_data is F, the filter is applied passively and the result is indicated by the fit variable output to the my_project_central.pts and my_project_perturbed.pts files.

  • The uncertainty interval is \(\pm \mathrm{sigma\_multiplier}\,\sigma_{i}\) where \(\sigma_{i}\) is, typically, the standard deviation of the analytical data entered in the *.imc file and sigma_multiplier is a scaling factor set by the sigma_multiplier option.

In mc_fit_plot:

  • The criterion can be used to filter *.pts file results that may contain models that do not fit all data within the specified uncertainty interval (i.e., files generated by mc_fit with must_fit_all_data set to F).

  • Because mc_fit_plot relies on passive application of the fit-all-data criterion by mc_fit, the mc_fit_plot option sigma_level must be set to the same value as the mc_fit option sigma_multiplier to ensure the imprecision interval used by mc_fit_plot to filter central model results corresponds to the uncertainty interval used by mc_fit to determine whether a model fits all data within uncertainty.

forward problem

The problem of predicting observations from a set of model parameters.

generic fluid solution model

A method of modeling mixing behavior in fluids that parallels conventional treatment of solid solutions ([Connolly2018]). In GFSM, the properties of each pure species are computed by an internal equation of state (EoS) specified by the user through perplex_option.dat and the non-ideality of the mixture is evaluated with an additional mixture EoS, typically the MRK. This contrasts with earlier treatments of complex fluids, Internal EoS, in which the fluid model was expressed entirely in terms of an internal algorithm ([Connolly1995]).

GFSM have the advantage of simplicity and flexibility over Internal EoS. Internal EoS are more precise and computationally efficient, but may involve assumptions about the components selected to represent the chemistry of a system and the chemistry itself (e.g., that the fluid is saturated in graphite).

GFSM are indicated by solution model code 39 in Perple_X solution model files. Internal EoS are indicated by solution model codes 0, 9, 20, 26, 40, 41, and 42.

GFSM

See generic fluid solution model.

imprecision interval

The interval of misfit values within which central models are considered indistinguishable from the best central model given the uncertainty in the input data. It is computed in mc_fit_plot and defined as

\([\ln(\mathrm{misfit})_{\rm BCM},\, \ln(\mathrm{misfit})_{\rm BCM} + \mathrm{sigma\_level}\, \epsilon_{\ln(\mathrm{misfit})}]\),

where BCM denotes the best central model, \(\epsilon_{\ln(\mathrm{misfit})}\) is the standard deviation of the perturbation analysis misfit scatter, and \(\mathrm{sigma\_level}\) specifies the desired coverage.

inverse problem

The problem of inferring model parameters from observations.

inversion parameters

The unknown parameters that are to be inferred from the known parameters of an inverse problem, i.e., the parameters on the left-hand side of Eq eq:inverse_problem. In the context of mc_fit, inversion parameters typically include temperature, pressure, and unmeasured compositional information.

misfit function

A quantitative measure of the difference between observed data and model predictions.

misfit imprecision

The uncertainty in the logarithm of the misfit arising from analytical and thermodynamic uncertainty. Models with misfit values within the imprecision interval of the best central model are considered indistinguishable given the uncertainty in the input data. It is quantified by the standard deviation of \(\ln(\mathrm{misfit})\) values obtained during perturbation analysis and is denoted \(\epsilon_{\ln(\mathrm{misfit})}\).

Monte Carlo sampling

A computational technique that uses random sampling to explore a parameter space. In mc_fit, Monte Carlo sampling is used to generate initial conditions for the central model analysis and to propagate analytical and thermodynamic uncertainties through the inversion process during the perturbation analysis.

MPP

In the inversion parameter space, the MPP is the centroid of the specified parameter ranges. If the ranges have been chosen rationally, the MPP represents the prior estimate of the most probable inversion parameter value. The MPP is also the average of the initial conditions generated for the Nelder-Mead optimizations of the central model analysis.

null phase

A phase with a composition that lies entirely within the mobile-component composition space (e.g., an H-O fluid in calculations as a function of hydrogen and oxygen chemical potentials). The null phase saturation surface bounds the equilibrium chemical potential coordinates for a system. Because null phases have no expression in the thermodynamic component space, the null phase saturation surface cannot be computed by conventional free energy minimization (e.g., vertex). Due to an algorithmic quirk, the null phase saturation surface can be computed with convex and the result superimposed on a phase diagram section computed by vertex.

observed-phase simplex

The phase simplex defined by the compositions of the observed phases.

perturbation analysis

The second stage of an mc_fit inversion. During this stage, the analytical and thermodynamic data are randomly perturbed within their uncertainties, and each optimization begins from the coordinates of the best central model. The resulting scatter quantifies data uncertainty and the misfit imprecision. The misfit imprecision is used to identify the retained central models from which the aleatoric uncertainty is estimated. Together, the data and aleatoric uncertainties determine the total uncertainty.

perturbed model

A forward model generated as part of the perturbation analysis, in which the analytical and thermodynamic data are perturbed within their estimated uncertainties and the optimization is initialized from the best central model. The population of perturbed models is used to estimate data uncertainty from the scatter of the inversion parameters and misfit imprecision from the scatter of the corresponding misfit values.

phase simplex

For a system with \(c\) components and an assemblage of \(p \le c\) equilibrium phases, the phase simplex is a \(p - 1\) dimensional simplex (the generalization of a 1-dimensional tie line) that connects the compositions of the \(p\) phases in the \(c - 1\) dimensional composition space. When \(p \gt c\), the phase simplex connects the compositions of the \(c\) phases that enclose the compositions of the remaining \(p - c\) phases. The phase simplex bounds the set of all possible positive linear combinations of the phase compositions. All bulk compositions that lie within the phase simplex have the same equilibrium phase assemblage, but the phase proportions vary as a function of the bulk composition.

In mc_fit, an incomplete bulk composition with \(c_\text{measured} \lt c\) components is sufficient to uniquely determine the phase proportions of the \(n = c_\text{measured}\) phases and the compositions of the remaining \(c - c_\text{measured}\) components. When \(n \lt p\), neither the bulk composition nor the phase proportions are unique unless modal data are provided.

potential variable

An intensive state variable (e.g., \(P, T, \mu_i, f_i, X_\text{CO2}^\text{fluid}\)) that is defined directly or indirectly by partial differentiation of a fundamental equation ([Callen1960], e.g., \(U(S,V,N)\)) with respect to an extensive variable (e.g., \(S, V,N\)) ([Tisza1961], [Callen1960], [Hillert1986], [Connolly1990]).

predicted-phase simplex

The phase simplex defined by the compositions of the phases predicted by mc_fit. In problems without modal or bulk compositional data, mc_fit tends to select a bulk composition that minimizes total misfit. When bulk compositional data are provided and \(p \ge c_\text{measured}\), successful optimizations uniquely determine the predicted phase proportions and bulk composition; additional constraints, such as modal data, have no effect on the predicted phase proportions or bulk composition. In this case, the solutions reproduce the measured part of the bulk composition exactly. When modal data are provided without bulk compositional data, successful optimizations uniquely determine the predicted phase proportions and bulk composition. In this case, the predicted phase proportions generally do not reproduce the observed phase proportions exactly because they are defined in the predicted-phase simplex, whereas the observed phase proportions lie in the observed-phase simplex. In problems with modal data and bulk compositional data, if \(p \lt c_\text{measured}\), successful solutions reproduce the measured bulk composition exactly and uniquely determine the predicted phase proportions and bulk composition.

total uncertainty

The uncertainty in the inversion parameters arising from both data uncertainty and aleatoric uncertainty. mc_fit_plot treats these contributions as independent, such that their variances and covariances are additive:

\[\sigma_{x,\mathrm{total}}^{2} = \sigma_{x,\mathrm{data}}^{2} + \sigma_{x,\mathrm{aleatoric}}^{2},\]
\[\sigma_{y,\mathrm{total}}^{2} = \sigma_{y,\mathrm{data}}^{2} + \sigma_{y,\mathrm{aleatoric}}^{2},\]
\[\mathrm{cov}_{xy,\mathrm{total}} = \mathrm{cov}_{xy,\mathrm{data}} + \mathrm{cov}_{xy,\mathrm{aleatoric}}.\]

Here, \(\sigma_x\) and \(\sigma_y\) are the standard deviations of inversion parameters \(x\) and \(y\), \(\mathrm{cov}_{xy}\) is their covariance, and the subscripts data, aleatoric, and total denote the corresponding uncertainty contributions.

This treatment assumes that the perturbation-analysis covariance varies negligibly over the imprecision interval. Although this assumption could be tested by repeating the perturbation analysis for multiple retained central models, such calculations are not performed because uncertainty propagation is expected to be invariant in the vicinity of the best central model.

For a single inversion parameter, the reported uncertainty interval is

\[x \pm \mathrm{sigma\_level}\,\sigma_{x,\mathrm{total}}.\]

Thus, for sigma_level = 1, the interval is the familiar \(1\,\sigma\) interval, corresponding to approximately 68% coverage for a normal distribution.

For two inversion parameters, the uncertainty ellipse is the locus of points satisfying

\[\begin{split}\begin{bmatrix} x-\bar{x} & y-\bar{y} \end{bmatrix} \mathbf{\Sigma}_{\mathrm{total}}^{-1} \begin{bmatrix} x-\bar{x} \\ y-\bar{y} \end{bmatrix} = \chi^2_2\!\left(p(\mathrm{sigma\_level})\right),\end{split}\]

where \(\bar{x}\) and \(\bar{y}\) are the best central model coordinates and \(\mathbf{\Sigma}_{\mathrm{total}}\) is the total covariance matrix,

\[\begin{split}\mathbf{\Sigma}_{\mathrm{total}} = \begin{bmatrix} \sigma_{x,\mathrm{total}}^2 & \mathrm{cov}_{xy,\mathrm{total}} \\ \mathrm{cov}_{xy,\mathrm{total}} & \sigma_{y,\mathrm{total}}^2 \end{bmatrix}.\end{split}\]

Here \(p(\mathrm{sigma\_level})\) is the probability enclosed by the 1D normal interval \(\pm \mathrm{sigma\_level}\,\sigma\) (e.g., \(p(\mathrm{sigma\_level}=1)\simeq0.68\), \(p(\mathrm{sigma\_level}=2)\simeq0.95\)), and \(\chi^2_2(p)\) is the chi-square value for two degrees of freedom that encloses probability \(p\).

try

A Nelder-Mead optimization to minimize the misfit between an observed phase assemblage and a thermodynamic model prediction as a function of the inversion parameters. Tries initiate from random initial guesses of the inversion parameters.

uncertainty

On input, mc_fit treats each analytical uncertainty as a sample standard deviation for the corresponding observed quantity. Within mc_fit, these standard deviations are multiplied by the factor specified by sigma_multiplier to define acceptance intervals for the observed quantities. Within mc_fit_plot, sigma_level specifies the coverage of the reported uncertainty intervals and covariance ellipses.

The uncertainty components reported by mc_fit_plot are represented by the symbols \(\sigma_{\rm data}\), \(\sigma_{\rm aleatoric}\), and \(\sigma_{\rm total}\). These symbols normally denote standard deviations. However, throughout this documentation the same notation is used for the corresponding reported uncertainty intervals. This convention is adopted for simplicity.

uncertainty ellipse

An ellipse representing the joint uncertainty of a pair of variables, defined by their covariance matrix. In mc_fit_plot, the ellipse is scaled according to sigma_level so that it has approximately the same coverage as the corresponding uncertainty intervals. Unlike uncertainty intervals, the ellipse accounts for covariance and represents both the magnitude and orientation of the joint uncertainty.

uncertainty interval

A symmetric interval about a reference value representing the uncertainty of an individual variable. In mc_fit_plot, the interval is scaled by sigma_level; for example, sigma_level = 1 corresponds approximately to 68% coverage, and sigma_level = 2 to 95% coverage. Uncertainty intervals provide a convenient means of expressing uncertainty in text, but, unlike uncertainty ellipses, represent only the marginal uncertainty of a variable.