Return Back
Logo ARPM
Logo ARPM
contact us
start here
  • Bootcamp
    • Overview
    • Reviews
    • Info/FAQs
    • Enroll
  • Certification
    • Overview
    • Reviews
    • Info/FAQs
    • Enroll
  • Lab/Study materials
    • Overview
    • Machine Learning
    • Quant Finance
    • Primers
    • What's new
    • Enroll
  • start here login
contact us
login
Introduction
About the ARPM Lab
Learning the ARPM Lab by topic
Learning the ARPM Lab by channel
Audience and prerequisites
Notation
Key tenets
Indices
Special characters
Distributions
Portfolio
Glossary
Symbols
A
B
C
D
E
F
G
H
I
J
K
L
M
N
O
P
Q
R
S
T
U
V
W
Y
Z
Bibliography
Data science
[1]Summary: “Data Science Map”
[1.1]Probabilistic framework
[1.2]Mean-covariance framework
[1.3]Linear models
[1.4]Machine learning
[1.5]Estimation
[1.6]Inference
[1.7]Sequential decisions
I. Probabilistic framework
[2]Environment: probabilistic
[2.1]Distributions
[2.1.1]Cumulative distribution function
[2.1.2]Probability density function
[2.1.3]Probability mass function
[2.1.4]Characteristic function
[2.1.5]Additional representations
[2.2]Visualization
[2.3]Transformations
[2.3.1]Information set
[2.3.2]Marginalization
[2.3.3]Push-forward
[2.3.4]Law of the unconscious statistician
[2.3.5]Change of measure
[2.4]Functionals
[2.4.1]Key definitions
[2.4.2]Mode
[2.4.3]Quantile family
[3]Structure: conditional independence
[3.1]Independence
[3.2]Conditioning
[3.2.1]Conditional distribution
[3.2.2]Geometrical interpretation
[3.3]Bayes theorem
[3.3.1]Model
[3.3.2]Joint
[3.3.3]Compound
[3.3.4]Posterior
[3.4]Reduced-form representation
[3.4.1]Independent component analysis
[3.4.2]Multivariate quantile function
[3.4.3]Innovation extraction
[3.4.4]Stochastic functions
[3.5]Conditional independence
[3.6]Conditional functionals
[3.6.1]Mode
[3.6.2]Mean
[3.6.3]Variance
[3.6.4]ANOVA
[3.6.5]Quantile
[4]Jointness: copulas
[4.1]Grades and inverse sampling
[4.2]Definition of copula
[4.2.1]Absolutely continuous distribution
[4.2.2]General distribution
[4.2.3]Copula density function
[4.3]Properties of copulas
[4.3.1]Independence
[4.3.2]Comonotonicity
[4.3.3]Invariance
[4.4]Elliptical copulas
[4.4.1]Normal copula
[4.4.2]t copula
[4.4.3]Scenarios from elliptical copulas
[4.5]Archimedean copulas
[4.5.1]Scenarios from Archimedean copulas
[4.6]Implementation
[4.6.1]Copula-marginal separation
[4.6.2]Copula-marginal combination
[5]Distance: information geometry
[5.1]Distributions geometry
[5.1.1]Tangent vector
[5.1.2]Fisher metric: length and volume
[5.1.3]Flatness and geodesics
[5.1.4]Duality: potentials and Legendre transformations
[5.1.5]Distance and divergence
[5.2]Exponential distributions geometry
[5.3]Scenario-probability distribution geometry
[6]Stochastic optimization: decision theory
[6.1]Statistical decision problems
[6.1.1]State
[6.1.2]Action
[6.1.3]Loss function
[6.1.4]Reward
[6.2]Stochastic dominance
[6.2.1]Strong dominance
[6.2.2]Weak dominance
[6.2.3]Higher order dominance
[6.2.4]Stochastic dominance failure
[6.3]Bayes risk
[6.3.1]Problem without inputs
[6.3.2]Problem with inputs
[6.3.3]Bayesian approach
[6.3.4]Frequentist approach
[6.3.5]A unified view
[6.3.6]Admissibility
[6.4]Minimax and other approaches
[6.4.1]Minimax
[6.4.2]Minimax versus Bayes
[6.4.3]Generalizations
[6.5]Causality
[6.5.1]Observational decision problem
[6.5.2]Interventional decision problem
[6.5.3]From observations to interventions
[6.5.4]Selection of interventional variables
[6.6]Scoring rules
[6.6.1]Propriety
[6.6.2]Notable scoring rules
[6.6.3]Loss-implied scoring rules
[6.6.4]Beyond scoring rules
[6.7]Points of interest, pitfalls, practical tips
[6.7.1]Randomized decisions
[6.7.2]Improper priors
[6.7.3]Propriety
[6.7.4]Expected value of perfect and sample information
[7]Stochastic programming
[7.1]Definitions
[7.2]General tricks
[7.3]Functional fit
[7.3.1]Machine learning as functional optimization
[7.3.2]Parametric functional optimization
[7.4]Linear basis
[7.4.1]Interactions/polynomials
[7.4.2]Orthogonal series
[7.4.3]Error minimization
[7.5]Trees
[7.5.1]CART
[7.5.2]Splines
[7.5.3]Voronoi diagrams
[7.5.4]Error minimization
[7.6]Neural networks
[7.6.1]Neurons
[7.6.2]Neural networks
[7.6.3]Projection pursuit
[7.6.4]Error minimization
[7.7]Gradient boosting
[7.8]RKHS/Kernel trick
II. Mean-covariance framework
[8]Environment: mean-covariance
[8.1]Mean-covariance classes
[8.1.1]Mean vector
[8.1.2]Covariance matrix
[8.1.3]Equivalence class
[8.1.4]Z-score
[8.2]Visualization
[8.2.1]Spectral decomposition
[8.2.2]Mean-covariance ellipsoid
[8.2.3]Principal component analysis
[8.3]Affine transformations
[8.3.1]Mean vector
[8.3.2]Covariance matrix
[8.3.3]Linearized information set
[8.3.4]Marginalization
[8.3.5]Z-score invariance
[8.3.6]Higher order moments
[8.4]Elicitability
[8.4.1]Individual versus joint
[8.4.2]Minimum volume ellipsoid
[8.4.3]Minimum z-score parameters
[9]Structure: partial uncorrelation
[9.1]Uncorrelation
[9.2]Linear projection
[9.2.1]Predicted mean-covariance class
[9.2.2]Geometrical interpretation
[9.3]Linear Bayes theorem
[9.3.1]Model
[9.3.2]Joint
[9.3.3]Compound
[9.3.4]Posterior
[9.4]Reduced-form representation
[9.4.1]Uncorrelated component analysis
[9.4.2]Linear quantile function
[9.4.3]Linear innovation extraction
[9.4.4]Linear stochastic functions
[9.5]Partial uncorrelation
[9.6]Connections with conditional independence
[9.6.1]The normal case
[9.6.2]Linear approximation
[10]Jointness: correlation
[10.1]Mean-covariance grades
[10.2]Definition of correlation
[10.3]Properties of correlation
[10.3.1]Uncorrelation
[10.3.2]Affine concordance
[10.3.3]Invariance
[10.4]Related definitions
[11]Stochastic optimization: mean-variance trade-off
[11.1]Problem without inputs
[11.1.1]Mean-variance trade-off
[11.1.2]Solutions
[11.1.3]Alternative formulations
[11.1.4]Special cases
[11.2]Problem with inputs
[11.2.1]Posterior risk
[11.2.2]Joint risk
[11.2.3]Sparsity
[11.2.4]Weak signal
[11.2.5]A unified result
[11.3]Points of interest
[11.3.1]Affine actions
[12]Location and dispersion
[12.1]Affine equivariance
[12.1.1]Key definitions
[12.1.2]Mode/modal dispersion
[12.1.3]Median/interquantile range
[12.1.4]Multivariate extensions
[12.2]Variational principles
[12.2.1]Key definitions
[12.2.2]Elicitable and quadrangle features
[12.2.3]Quantile family
[12.2.4]Mean absolute deviation family
[12.2.5]Median absolute deviation
[12.2.6]Other locations/dispersions
[12.2.7]Multivariate extensions
[12.3]Points of interest, pitfalls, practical tips
[12.3.1]Visualization via uncertainty bands
[13]Dependence and concordance
[13.1]Measures of dependence
[13.1.1]Schweizer-Wolff measure
[13.1.2]Mutual information
[13.2]Measures of concordance
[13.2.1]Kendall’s tau
[13.2.2]Spearman’s rho
[13.3]Correlation as measure of dependence/concordance
[13.4]Points of interest, pitfalls, practical tips
[13.4.1]Schweizer and Wolff measure via simulations
III. Linear models
[14]Linear factor models
[14.1]Definitions
[14.1.1]The r-squared
[14.1.2]Dominant-residual models
[14.1.3]Systematic-idiosyncratic models
[14.1.4]Estimation
[14.2]Linear least squares regression models
[14.2.1]Definition
[14.2.2]Solution: factor loadings
[14.2.3]Prediction and fit
[14.2.4]Residuals features
[14.2.5]Natural scatter specification
[14.2.6]Estimation
[14.3]Principal component models
[14.3.1]Definition
[14.3.2]Solution: factor loadings and constructed factors
[14.3.3]Identification issues
[14.3.4]Prediction and fit
[14.3.5]Residuals features
[14.3.6]Natural scatter specification
[14.3.7]Estimation
[14.4]Factor analysis models
[14.4.1]Definition
[14.4.2]Solution: factor loadings and idiosyncratic variances
[14.4.3]Exact principal component with isotropic variances
[14.4.4]Identification issues
[14.4.5]Factor scores
[14.4.6]Prediction and fit
[14.4.7]Residuals features
[14.4.8]Natural scatter specification
[14.4.9]Estimation
[14.5]Cross-sectional models
[14.5.1]Definition
[14.5.2]Solution: factor-construction matrix
[14.5.3]Prediction and fit
[14.5.4]Residuals features
[14.5.5]Natural scatter specification
[14.5.6]Systematic-idiosyncratic assumption
[14.5.7]Estimation
[14.6]Points of interest, pitfalls, practical tips
[14.6.1]LFM’s are not a regression on past data(1) [⋆⋆]
[14.6.2]LFM’s are not about returns(2-3) [⋆⋆]
[14.6.3]LFM’s are not about stocks(4) [⋆⋆]
[14.6.4]LFM’s “factors”are not “factors returns”(5) [⋆]
[14.6.5]LFM’s are not systematic-idiosyncratic(6-7) [⋆⋆⋆]
[14.6.6]LFM’s are not horizon-independent(8) [⋆⋆]
[14.6.7]LFM’s are not a dimension reduction technique(9) [⋆]
[14.6.8]LFM’s are not APT and CAPM(10-11-12) [⋆⋆⋆]
[14.6.9]Factor analysis LFM’s are not idiosyncratic(14) [⋆⋆]
[14.6.10]LFM’s do not extract premia-generating factors(16)[⋆⋆]
[14.6.11]LFM’s are not always necessary(13-15-17-18) [⋆⋆⋆]
[14.6.12]Affine versus linear formulation
[14.6.13]Linear regression: a success story
[14.6.14]Principal factors are not principal components
[14.6.15]Performance of regression versus principal component
[14.6.16]Conditional principal component
[14.6.17]More general constraints
[15]Implicit dominant-residual models
[16]Structural equation models
IV. Machine learning
[17]Foundations
[17.1]Approaches to machine learning
[17.1.1]Supervised learning
[17.1.2]Unsupervised learning
[17.1.3]Reinforcement learning
[17.1.4]Linear factor models roots
[17.2]Prediction
[17.2.1]Point prediction
[17.2.2]Probabilistic prediction
[17.2.3]Point/probabilistic connections
[17.3]Learning and inference
[17.3.1]Learning
[17.3.2]Inference
[17.4]Points of interest, pitfalls, practical tips
[17.4.1]Marginalization and mode computation
[18]Supervised learning: regression
[18.1]Point least squares regression
[18.1.1]Error
[18.1.2]Theoretical optimum
[18.1.3]Optimization in practice
[18.1.4]Linear basis
[18.1.5]ANOVA
[18.1.6]Trees
[18.1.7]Neural networks
[18.1.8]Gradient boosting
[18.1.9]Kernel trick
[18.1.10]Geometrical interpretation
[18.2]Point non-least squares regression
[18.2.1]Error
[18.2.2]Theoretical optimum
[18.2.3]Optimization in practice
[18.2.4]Linear basis
[18.2.5]Advanced methods
[18.2.6]Generalized point regression
[18.3]Probabilistic regression
[18.3.1]Error
[18.3.2]Theoretical optimum
[18.3.3]Target parameters
[18.3.4]Optimization in practice
[18.3.5]Linear regression
[18.3.6]Generalized linear models
[18.3.7]Alternative generalizations
[18.4]Points of interest, pitfalls, practical tips
[18.4.1]Alternative errors
[18.4.2]Exponential tilting
[19]Supervised learning: classification
[19.1]Point binary classification
[19.1.1]Joint distribution
[19.1.2]Error
[19.1.3]Theoretical optimum
[19.1.4]Receiver operating characteristic (ROC)
[19.1.5]Optimization in practice
[19.1.6]Perceptron
[19.1.7]Support vector machines
[19.1.8]Fisher discriminant analysis
[19.2]Point multinomial classification
[19.2.1]Joint distribution
[19.2.2]Error
[19.2.3]Theoretical optimum
[19.2.4]Classification via discriminants
[19.2.5]Optimization in practice
[19.2.6]Leveraging binary classifiers
[19.3]Probabilistic classification
[19.3.1]Error
[19.3.2]Theoretical optimum
[19.3.3]Alternative approaches
[19.3.4]Target parameters
[19.3.5]Optimization in practice
[19.3.6]Logistic regression
[19.3.7]Naive Bayes classifiers
[19.3.8]Probit regression
[19.3.9]Neural networks
[19.3.10]Trees
[19.3.11]Gradient boosting
[19.4]Points of interest, pitfalls, practical tips
[19.4.1]Binary classification: alternative errors
[19.4.2]Linear regression for classification
[20]Unsupervised learning: autoencoders
[20.1]Least squares autoencoders
[20.1.1]Minimum torsion variables
[20.1.2]k-means clustering
[20.1.3]Kernel trick
[21]Unsupervised learning: graphical models
[21.1]Graphs
[21.2]Definitions
[21.3]Probabilistic factor analysis
[21.3.1]Solution: factor loadings and idiosyncratic variances
[21.3.2]Maximum likelihood factorization algorithm
[21.3.3]Special case: isotropic variances
[21.3.4]Identification issues
[21.3.5]Inference
[21.4]Mixture models
[21.4.1]Gaussian mixture models
[21.4.2]EM algorithm for Gaussian mixture models
[21.4.3]Inference
[21.4.4]Mixture of experts
[21.4.5]General mixtures models
[21.5]Naive Bayes models
[21.6]Bayes networks
[21.7]Markov networks
[22]Optimal transport
[22.1]General case
[22.1.1]Transport maps
[22.1.2]Monge problem
[22.1.3]Couplings
[22.1.4]Kantorovich problem
[22.1.5]Dual problem
[22.1.6]Wasserstein distance
[22.2]Categorical distributions
[22.2.1]Transport maps
[22.2.2]Monge problem
[22.2.3]Couplings
[22.2.4]Kantorovich problem
[22.2.5]Dual problem
[22.2.6]Wasserstein distance
[22.3]Histograms
[22.3.1]Transport maps
[22.3.2]Couplings
[22.3.3]Kantorovich problem
[22.3.4]Wasserstein distance
[22.4]Squared 2-norm loss
[22.4.1]Monge problem
[22.4.2]Kantorovich problem
[22.4.3]Dual problem
[22.4.4]Wasserstein distance
[22.4.5]Multivariate quantile function
[22.4.6]Polar decomposition
[22.4.7]Voronoi cells
[22.5]Linear optimal transport
[22.5.1]Linear transport maps
[22.5.2]Linear Monge problem
[22.5.3]Linear couplings
[22.5.4]Linear Kantorovich problem
[22.5.5]Dual Kantorovich problem
[22.5.6]Wasserstein distance
[22.5.7]Polar decomposition
[22.6]Earth mover problem
[22.6.1]General distance
[22.6.2]Histograms, q-norm
[22.6.3]Histograms, 0-1 distance
V. Estimation
[23]Estimation primer
[23.1]Flexible probabilities
[23.1.1]Exponential decay and time conditioning
[23.1.2]Kernels and state conditioning
[23.1.3]Joint state and time conditioning
[23.1.4]Statistical power of flexible probabilities
[23.2]Historical estimation
[23.2.1]Canonical historical estimation
[23.2.2]Generalization to flexible probabilities
[23.2.3]Extracting properties
[23.2.4]Exponential moving moments and statistics
[23.3]Kernel estimation
[23.3.1]Canonical kernel estimation
[23.3.2]Generalization to flexible probabilities
[23.4]Maximum likelihood estimation
[23.4.1]Maximum likelihood principle
[23.4.2]Canonical maximum likelihood for i.i.d. variables
[23.4.3]Generalization to flexible probabilities
[23.4.4]Extracting properties
[23.4.5]Exponential family (i.i.d.) assumption
[23.4.6]Generalization to non-i.i.d. observable processes
[23.5]Hidden variables and missing data
[23.5.1]Latent variables
[23.5.2]EM algorithm
[23.5.3]IID latent processes
[23.5.4]Markov state-space processes
[23.5.5]Networks
[23.5.6]Missing data
[23.6]Generalized method of moments
[23.6.1]Canonical method of moments
[23.6.2]Generalization to flexible probabilities
[23.6.3]Generalized method of moments - exact specification
[23.6.4]Generalized method of moments - over specification
[23.7]Robustness
[23.7.1]Local robustness
[23.7.2]Global robustness
[23.8]Bayesian estimation
[23.8.1]Estimation
[23.8.2]Prediction
[23.8.3]Analytical solutions
[23.8.4]Exponential family (i.i.d.) assumption
[23.9]Points of interest, pitfalls, practical tips
[23.9.1]Unconditional estimation
[23.9.2]Outlier detection
[23.9.3]Backward/forward exponential decay
[24]Estimation: location and dispersion
[24.1]Historical
[24.1.1]HFP mean, covariance and correlation
[24.1.2]HFP mean-covariance ellipsoid
[24.2]Maximum likelihood
[24.2.1]Normal assumption
[24.2.2]Student t assumption
[24.3]Missing data
[24.3.1]Randomly missing data
[24.3.2]Times series of different length
[24.4]Robustness
[24.4.1]High breakdown point with flexible probabilities estimators: theory
[24.4.2]High breakdown point with flexible probabilities estimators: practice
[24.5]Bayesian
[24.5.1]Model and sample estimators
[24.5.2]Normal-inverse-Wishart prior distribution
[24.5.3]Normal-inverse-Wishart posterior distribution
[24.5.4]Student t predictive distribution
[24.5.5]Classical equivalent, uncertainty and shrinkage
[24.6]Shrinkage
[24.6.1]Mean shrinkage: James-Stein
[24.6.2]Covariance shrinkage: Ledoit-Wolf
[24.6.3]Correlation shrinkage: random matrix theory
[24.6.4]Covariance shrinkage: sparse eigenvector rotations
[24.6.5]Covariance shrinkage: glasso
[24.6.6]Covariance shrinkage: factor analysis
[24.7]Mixed approach
[24.8]Frequentist risk: analytical results
[24.8.1]Sample mean/covariance
[25]Estimation: regression loadings
[25.1]Historical
[25.2]Maximum likelihood
[25.2.1]The model
[25.2.2]Normal assumption
[25.2.3]Student t assumption
[25.3]Bayesian
[25.3.1]The model
[25.3.2]Conditional likelihood and sample estimators
[25.3.3]Normal-inverse-Wishart prior distribution
[25.3.4]Normal-inverse-Wishart posterior distribution
[25.3.5]Student t predictive distribution
[25.3.6]Classical equivalent
[25.3.7]Uncertainty
[25.4]Regularization: factors selection
[25.4.1]Stepwise regression selection
[25.4.2]Lasso regression
[25.4.3]Ridge regression
[25.5]Mixed approach
[26]Random matrix theory
[26.1]Random matrix ensembles
[26.1.1]Eigenvalues
[26.1.2]Eigenvectors
[26.2]Empirical spectral distribution
[26.2.1]Probability density functions
[26.2.2]Cumulative distribution and quantile
[26.2.3]Moments
[26.2.4]Stieltjes transform
[26.2.5]Resolvent
[26.2.6]Replica
[26.2.7]Orthogonal transformations
[26.3]Infinite matrix limit
[26.3.1]Representations
[26.3.2]Deterministic convergence
[26.3.3]Stochastic convergence
[26.4]Boltzmann ensembles
[26.4.1]The ensembles
[26.4.2]Eigenvalues
[26.4.3]Eigenvectors
[26.4.4]Infinite matrix limit
[26.5]Wigner ensemble
[26.5.1]The ensemble
[26.5.2]Gaussian orthogonal ensemble
[26.5.3]Beyond Gaussian orthogonal
[26.6]Marchenko-Pastur ensemble
[26.6.1]The ensemble
[26.6.2]Addressing singularity
[26.6.3]Infinite matrix limit
[26.6.4]Wishart orthogonal ensemble
[26.6.5]Beyond Wishart orthogonal ensemble
[26.7]Free probability
[26.7.1]Scalars
[26.7.2]Matrices
[26.7.3]Infinite matrix limit
[26.7.4]Transforms
[26.7.5]Sums
[26.7.6]Products
[26.8]Dense covariances
[26.8.1]Samples with dense covariance
[26.8.2]Sample covariance revisited
[26.8.3]Limiting spectral density
[26.8.4]Application
[26.9]Spiked covariances
[26.9.1]Samples from linear factor models
[26.9.2]Sample covariance revisited
[26.9.3]Limiting spectral density
[26.9.4]Characteristic polynomial
[26.9.5]Free probability approach
[26.9.6]Outlier of the full covariance matrix
[27]Estimation theory: classical framework
[27.1]Definitions
[27.2]Bayesian
[27.3]Frequentist
[27.3.1]Risk
[27.3.2]Bias versus variance
[27.3.3]Minimax estimator
[27.4]Points of interest, pitfalls, practical tips
[27.4.1]Sample quantiles (order statistics)
[28]Hypothesis testing
[28.1]Single binary testing
[28.1.1]Tests
[28.1.2]P-value test
[28.2]Multiple binary testing
[28.3]Hypothesis testing for invariants
[28.3.1]Univariate testing: the z-statistic
[28.3.2]Multivariate testing: the Hotelling statistic
[29]Estimation theory: elicitable framework
[29.1]Definitions
[29.1.1]Predictive decision problem
[29.1.2]Estimative decision problem
[29.1.3]Sub-optimal two-step estimation
[29.1.4]New goal: oracle decision
[29.1.5]Optimal one-step estimation
[29.1.6]Conditional independence
[29.1.7]A unified view
[29.1.8]Conclusions
[29.2]Bayesian
[29.3]Frequentist
[29.3.1]Excess risk
[29.3.2]Empirical risk minimization
[29.3.3]Bias versus variance
[29.3.4]Approximation versus estimation
[29.3.5]Generalizations
[30]Estimation risk mitigation
[30.1]Bayesian
[30.2]Frequentist
[30.2.1]Ensemble
[30.2.2]Regularization
[30.2.3]Cross-validation
[30.2.4]Information criteria
[30.2.5]Best estimator
[30.3]Ensemble learning
[30.3.1]Bagging
[30.3.2]Flexible probabilities as random-variables
[30.3.3]Flexible probabilities through conditioning
[30.3.4]Ensemble weighting
[30.4]Regularization
[30.4.1]Stepwise features selection
[30.4.2]Ridge, lasso, elastic nets
[30.4.3]Glasso
[30.4.4]Categorical factors selection
[30.4.5]Bayesian prior
[30.4.6]Sparse principal component
[30.5]Cross-validation
[30.5.1]Background
[30.5.2]Estimation: in-sample error
[30.5.3]Testing: out-of-sample error
[30.5.4]Best estimator
[30.5.5]Cases of interest
[30.6]Information criteria and asymptotic theory
[30.7]Quest for invariance
[31]Invariance tests
[31.1]Simple tests
[31.2]Refinements and pitfalls
[31.2.1]Circle-like covariance (not data)
[31.2.2]Stronger tests based on copulas
VI. Inference
[32]Black-Litterman
[32.1]Prior distribution
[32.1.1]Performance model
[32.1.2]Prior distribution of expected returns
[32.1.3]Prior predictive performance distribution
[32.2]Active views
[32.2.1]Active views model
[32.2.2]Active views statement
[32.2.3]Posterior distribution of the expected returns
[32.2.4]Posterior predictive distribution
[32.3]Limit cases and generalizations
[32.3.1]High confidence in prior
[32.3.2]Low confidence in views
[32.3.3]High confidence in views
[32.3.4]Generalizations
[32.3.5]From linear returns to risk drivers
[32.3.6]From stock-like to generic asset classes
[32.3.7]From normal to non-normal markets
[32.3.8]From linear equality views to partial flexible views
[33]Generalized probabilistic inference
[33.1]Views processing: minimum relative entropy
[33.1.1]Base distribution and view variables
[33.1.2]Point views
[33.1.3]Distributional views
[33.1.4]Partial views
[33.1.5]Partial views on generalized expectations
[33.1.6]Sanity check
[33.1.7]Confidence
[33.1.8]Relationship with Bayesian updating
[33.2]Analytical implementation
[33.2.1]Base distribution
[33.2.2]Views
[33.2.3]Sanity check
[33.2.4]Updated distribution
[33.2.5]Confidence
[33.2.6]Relevant special cases
[33.3]Flexible probabilities implementation
[33.3.1]Base distribution
[33.3.2]Views
[33.3.3]Sanity check
[33.3.4]Updated distribution
[33.3.5]Confidence
[33.4]Factor-based implementations
[33.5]Copula opinion pooling
[33.5.1]Base distribution
[33.5.2]Views
[33.5.3]Updated distribution
[33.5.4]Confidence
[33.5.5]The algorithm
[33.6]Generalized shrinkage
[33.6.1]Intuition
[33.6.2]Classical shrinkage
[33.6.3]Bayesian updating
[33.6.4]Minimum relative entropy
[33.6.5]Shrinkage
[33.6.6]Regularization
[34]Inference via Monte Carlo and variational techniques
[34.1]Inference via Monte Carlo
[34.1.1]Metropolis-Hastings
[34.2]Inference and learning via variational techniques
[34.2.1]IM projection
[34.2.2]Inference
[34.2.3]Learning
[34.2.4]Analytical solution: exponential family
[34.2.5]Variational solution
[34.2.6]EM algorithm in population
[34.2.7]Dimension reduction
VII. Sequential decisions
[35]Stochastic processes environment
[35.1]Definitions
[35.1.1]Stochastic processes
[35.1.2]Paths
[35.1.3]Probabilistic specification
[35.1.4]Mean-covariance kernels
[35.2]Relevant properties
[35.2.1]Strong stationarity
[35.2.2]Covariance stationarity
[35.2.3]Ergodicity
[35.3]Prediction
[35.3.1]Probabilistic prediction
[35.3.2]Krieging
[35.3.3]Financial applications
[35.4]Points of interest
[35.4.1]Autocorrelation kernel
[35.4.2]Mean-covariance random fields
[35.4.3]Granger causality
[35.4.4]Linear decomposition
[35.4.5]Conditional expectation as best prediction
[35.4.6]General representation
[36]Random walk
[36.1]Strong white noise
[36.2]Discrete time random walk
[36.2.1]Definitions
[36.2.2]Relevant cases
[36.2.3]Forecast
[36.3]Levy processes
[36.3.1]Infinite divisibility
[36.3.2]Continuous state: Brownian diffusion
[36.3.3]Discrete state: Poisson jumps
[36.3.4]Notable Levy processes
[36.3.5]Levy-Khintchine representation
[36.3.6]Subordination
[36.3.7]Fourier algorithm for non- divisible processes
[36.4]Square-root rule and generalizations
[36.4.1]Thin-tailed random walk
[36.4.2]Thick-tailed random walk
[36.4.3]Multivariate random walk
[36.4.4]General processes
[36.5]Martingales
[37]Autoregressive processes
[37.1]Weak white noise
[37.1.1]Definition
[37.1.2]Relevant cases
[37.2]Autoregression of order one
[37.2.1]Definitions
[37.2.2]Stationarity
[37.2.3]Forecast
[37.3]Vector autoregression of order one
[37.3.1]Definitions
[37.3.2]Relevant cases
[37.3.3]Stationarity
[37.3.4]Cointegration
[37.3.5]Estimation
[37.3.6]Prediction
[37.4]Linear state-space models
[37.4.1]Definitions
[37.4.2]Relevant cases
[37.4.3]Stationarity
[37.4.4]Estimation
[37.4.5]Prediction - Kalman filter
[37.5]Ornstein-Uhlenbeck process
[37.5.1]Forecast and conditional distribution of OU
[37.5.2]Stationarity and unconditional distribution
[37.6]Multivariate Ornstein-Uhlenbeck
[37.6.1]Definitions
[37.6.2]Forecast and conditional distribution of MVOU
[37.6.3]Stationarity and unconditional distribution of MVOU
[37.6.4]Geometrical interpretation∗
[37.6.5]Cointegrated Ornstein-Uhlenbeck
[37.6.6]Relationship between (V)AR and (MV)OU
[37.7]Orthogonal increment processes
[38]Covariance stationary theory
[38.1]Spectral representation
[38.1.1]Spectral theorem - intuition
[38.1.2]Spectral theorem - formal statement
[38.1.3]Cramer decomposition - intuition
[38.1.4]Cramer decomposition - formal statement
[38.1.5]Application: identification
[38.2]Filtering
[38.2.1]Intuition
[38.2.2]Formal definitions
[38.2.3]autocovariance function
[38.2.4]Spectrum
[38.2.5]Affine equivariance
[38.2.6]Composition, inversion
[38.2.7]Causality
[38.2.8]Time domain filters
[38.2.9]Frequency domain filters
[38.3]Wold representation
[38.3.1]Intuition
[38.3.2]Formal statement
[38.3.3]Relationship with spectral analysis
[38.3.4]Computation of Wold components
[38.4]Dynamic factor models
[38.4.1]Dynamic regression
[38.4.2]Dynamic principal components
[39]Other mean-covariance stochastic models
[39.1](V)ARMA processes
[39.1.1]Definitions
[39.1.2]Stationarity
[39.1.3]Invertible (V)ARMA
[39.1.4]Forecast
[39.2]Integrated processes
[39.2.1]Integrated of order zero process
[39.2.2]Integer integration: ARIMA
[39.2.3]Fractional integration: fractional white noise
[39.3]Fractional Brownian motion
[39.4]Harmonic processes
[39.4.1]Definitions
[39.4.2]Relevant cases
[39.4.3]Mean and autocovariance
[39.4.4]Spectral density
[39.4.5]General basis as AR(2) limit
[39.4.6]Periodic harmonics as lagged AR(1) limit
[39.4.7]Multivariate harmonics
[39.4.8]Harmonics forecast
[39.5]Polynomial trend processes
[39.5.1]Definitions
[39.5.2]Stochastic approximations of deterministic trends
[39.5.3]Forecast
[40]Wiener-Kolmogorov filter
[40.1]From regression to filter
[40.2]Endogenous Wiener-Kolmogorov filter
[40.3]Exogenous Wiener-Kolmogorov filter
[41]Relevant probabilistic stochastic models
[41.1]Markov processes
[41.1.1]Theory
[41.1.2]Relevant cases
[41.2]Markov chains
[41.2.1]Time-homogeneous Markov chains
[41.2.2]Time-inhomogeneous Markov chains
[41.2.3]Multivariate Markov chain
[41.2.4]Continuous time-homogeneous Markov chain
[41.2.5]Continuous time-inhomogeneous Markov chains
[41.2.6]Stationarity and unconditional distributions
[41.3]State space processes
[41.3.1]Theory
[41.3.2]Probabilistic linear state-space models
[41.3.3]Hidden Markov models
[41.3.4]Hidden Markov VAR(1) models
[41.4]GARCH(1,1) process
[41.5]Stochastic volatility models
[41.5.1]State-space stochastic volatility
[41.5.2]Discrete time Heston model
[41.5.3]Hybrid models
[41.5.4]Continuous time Heston model
[41.5.5]Time changed Brownian motion
[41.5.6]Connection between time-changed Brownian motion and stochastic volatility
[41.6]Points of interest
[41.6.1]Probabilistic graphical models
[41.6.2]Markov property for random fields
[42]State-space forecasting
[42.1]Markov processes forecast
[42.1.1]Monte Carlo
[42.1.2]Historical bootstrapping
[42.1.3]Arbitrary monitoring times
[42.2]State-space processes forecast
[42.3]Points of interest
[42.3.1]Probabilistic forecast for general models
[42.3.2]Scenario projection enhancements by probability twisting
[42.3.3]Hybrid Monte Carlo-historical
VIII. Data science toolbox
[43]Linear algebra
[43.1]Vector spaces
[43.1.1]Vector operations
[43.1.2]Basis and coordinates
[43.1.3]Vector subspaces
[43.2]Linear transformations
[43.2.1]Matrix representation
[43.2.2]Composition
[43.2.3]Invertibility
[43.3]Inner product spaces
[43.3.1]Symmetry
[43.3.2]Positivity
[43.3.3]Length, distance and angle
[43.3.4]Orthogonal projection
[43.3.5]Best prediction
[43.3.6]Rotations
[43.4]Metric and normed spaces
[43.4.1]Norm
[43.4.2]Distance
[43.4.3]Divergence
[43.4.4]Geodesics
[43.5]Spectral decomposition
[43.5.1]Eigenvalues and eigenvectors
[43.5.2]Square matrix spectral decomposition
[43.5.3]Spectral theorem
[43.5.4]Singular value decomposition
[43.6]Matrix transpose-square-root
[43.6.1]Gramian
[43.6.2]Transpose-square-roots
[43.6.3]Orthonormalization
[43.6.4]Spectrum/principal components
[43.6.5]Cholesky/Gram-Schmidt
[43.6.6]Riccati/minimum torsion
[43.7]Matrix operations
[43.7.1]The vector space of matrices
[43.7.2]Key operations
[43.7.3]Pseudo-inverse
[43.7.4]Useful identities
[43.8]Matrix polynomials
[43.8.1]Matrix polynomials factorization
[43.8.2]Matrix polynomial inversion
[43.9]Pitfalls and points of interest
[43.9.1]Multiplicities
[44]Calculus
[44.1]Differentiation
[44.1.1]Univariate functions
[44.1.2]Multivariate functions
[44.1.3]Matrix-variate functions
[44.2]Taylor expansion
[44.2.1]Univariate functions
[44.2.2]Multivariate functions
[44.3]Integration
[44.3.1]Partitions and measurability
[44.3.2]Univariate integration
[44.3.3]Fundamental theorem of calculus
[44.3.4]Multivariate integration
[44.4]Monotone functions
[44.4.1]Univariate monotonicity
[44.4.2]Entrywise monotonicity
[44.4.3]Monotone maps
[44.5]Convexity
[44.5.1]Univariate convexity
[44.5.2]Multivariate convexity
[45]Optimization
[45.1]Fundamental concepts
[45.1.1]The optimization problem
[45.1.2]Local minimum
[45.2]Smooth programming
[45.2.1]First and second order criteria
[45.2.2]Lagrange multipliers
[45.2.3]Gradient descent
[45.2.4]Newton’s method
[45.3]Convex programming
[45.3.1]The general problem
[45.3.2]Linear programming
[45.3.3]Quadratic programming
[45.3.4]Second-order cone programming
[45.3.5]Semidefinite programming
[45.3.6]Conic programming
[45.4]Quadratic regularization
[45.4.1]Ridge regularization
[45.4.2]Lasso regularization
[45.4.3]Elastic net regularization
[45.5]Selection problems
[45.5.1]Problem statement
[45.5.2]General solution
[45.5.3]Combinatorial heuristics
[45.5.4]Elastic net heuristics
[45.6]Equivalent optimization problems
[45.6.1]Invertible function of the objective
[45.6.2]Epigraph form
[45.6.3]Slack variables
[46]Functional analysis
[46.1]Measure theory
[46.1.1]Domains
[46.1.2]Measures
[46.1.3]Lebesgue’s decomposition
[46.2]Functional algebra
[46.2.1]Function spaces
[46.2.2]Linear operators
[46.2.3]Kernel representation
[46.2.4]Eigenvalues and eigenfunctions
[46.3]L2 spaces
[46.3.1]Inner product
[46.3.2]Dirac delta
[46.3.3]Riesz representation theorem
[46.3.4]Unitary operators
[46.3.5]Lp geometry
[46.4]Fourier transform
[46.4.1]Intuition
[46.4.2]Toeplitz structure
[46.4.3]General transform and convolution
[46.4.4]Fourier integral transform
[46.4.5]Discrete time Fourier transform
[46.4.6]Fourier series
[46.4.7]Discrete Fourier transform
[46.5]Spectral theorem
[46.5.1]Motivation
[46.5.2]Mercer kernels
[46.5.3]Matrix-valued kernels
[46.6]Bochner’s theorem
[46.6.1]Spectral representation
[46.6.2]Power spectrum
[46.6.3]Matrix-valued kernels
[46.7]Mercer’s theorem
[46.7.1]Spectral representation
[46.7.2]Reproducing kernel Hilbert spaces
[46.7.3]Matrix-valued kernels
[46.8]Functional calculus
[46.8.1]Gateaux derivative
[46.8.2]Fréchet derivative
[46.8.3]Second order derivative
[47]Discrete mathematics
[47.1]Discrete derivatives
[47.1.1]Univariate derivatives
[47.1.2]Multivariate derivatives
[47.2]Combinatorial programming
[47.2.1]Brute force search
[47.2.2]Naive selection
[47.2.3]Stepwise forward selection
[47.2.4]Stepwise backward elimination
[47.2.5]2-step forward heuristic
[47.2.6]General combinatorial programming heuristics
[48]Abstract probability
[48.1]Key concepts
[48.1.1]Probability space
[48.1.2]Random variable
[48.1.3]Random fields
[48.1.4]Expectation
[48.1.5]Radon-Nikodym derivative
[48.1.6]Abstract distributions
[48.1.7]Conditional probability
[48.2]L2 spaces of random variables
[48.2.1]Inner product
[48.2.2]Length, distance and angle
[48.2.3]Visualization
[48.2.4]Geometry of random vectors
[48.2.5]Projection
[48.2.6]Covariance (improper) inner product
[48.3]Abstract conditional expectation
[48.3.1]Partitions of the sample space
[48.3.2]Probability conditional on a partition
[48.3.3]Discretization of random variables
[48.3.4]Conditional discretization of random variables
[48.3.5]Abstract Bayes theorem
[48.4]Abstract stochastic processes
[48.4.1]Filtrations
[48.4.2]Iterated expectations
[48.4.3]Adapted processes
[48.4.4]Martingales
[48.4.5]Approximations of processes
[49]Notable distributions
[49.1]Normal
[49.1.1]Pdf, cdf and characteristic function
[49.1.2]Moments
[49.1.3]Conditional distribution
[49.1.4]Stochastic representations
[49.1.5]Affine equivariance
[49.1.6]Matrix-normal
[49.1.7]Gaussian random fields
[49.2]Lognormal
[49.2.1]Pdf, cdf and characteristic function
[49.2.2]Moments
[49.2.3]Conditional distribution
[49.2.4]Shifted lognormal
[49.3]Quadratic normal
[49.3.1]Chi-squared
[49.3.2]Gamma
[49.3.3]Generalized chi-squared
[49.3.4]Wishart
[49.3.5]Inverse-Wishart
[49.4]Elliptical distributions
[49.4.1]Fundamental concepts
[49.4.2]Student t
[49.4.3]Cauchy
[49.4.4]Uniform inside the ellipsoid
[49.4.5]Uniform on the ellipsoid
[49.4.6]Affine equivariance
[49.4.7]Stochastic representations
[49.4.8]Generation of elliptical scenarios
[49.4.9]Scenario generation with dimension reduction
[49.5]Scenario-probability
[49.5.1]Types of scenario-probability distributions
[49.5.2]Probability mass and density function
[49.5.3]Transformations and generalized expectations
[49.5.4]Cumulative distribution function
[49.5.5]Quantile
[49.5.6]Moments and other statistical features
[49.6]Categorical
[49.6.1]Discriminant variables
[49.6.2]Probabilities parametrization
[49.7]Exponential family
[49.7.1]Normal
[49.7.2]Categorical
[49.8]Mixtures
[49.8.1]Binary case
[49.8.2]Multinomial case
[49.9]Stable, additive and infinitely divisible
[49.9.1]Stable
[49.9.2]Additive
[49.9.3]Infinitely divisible
[49.10]Moment-matching scenarios
[49.10.1]Twisting scenarios
[49.10.2]Twisting probabilities
Quantitative finance
[50]Summary: “Quantitative Finance Checklist”
[50.1]Financial engineering
[50.2]Risk management
[50.3]Portfolio management
[50.4]P versus Q
IX. Financial engineering
[51]Step 1: Valuation
[51a]Step 1a: Linear pricing theory - core
[51a.1]Fundamental axioms
[51a.1.1]Law of one price
[51a.1.2]Linearity
[51a.1.3]Absence of arbitrage
[51a.1.4]Relationships among fundamental axioms
[51a.2]Fundamental theorem of asset pricing
[51a.2.1]Linear pricing equation
[51a.2.2]Numeraire
[51a.2.3]Identification issues
[51a.3]Risk-neutral pricing
[51a.3.1]Discrete-time rebalancing
[51a.3.2]No rebalancing: forward measure
[51a.3.3]Continuous rebalancing limit
[51a.4]Capital asset pricing framework
[51a.4.1]Maximum Sharpe ratio portfolio
[51a.4.2]Security market line
[51a.4.3]Connections to CAPM and linear factor models
[51a.5]Covariance principle
[51a.5.1]Risk premium and equivalence with the security market line
[51a.5.2]Credit
[51a.5.3]Buhlmann exponential tilting
[51b]Step 1b: Linear pricing theory - further assumptions
[51b.1]Completeness
[51b.1.1]General statement
[51b.1.2]Arrow-Debreu securities
[51b.1.3]European options
[51b.2]Equilibrium: capital asset pricing model
[51b.3]Arbitrage pricing theory
[51b.3.1]Standard derivation: linear factor model for instruments
[51b.4]Intertemporal consistency
[51b.4.1]The framework
[51b.4.2]Intertemporal linear pricing equation
[51b.4.3]Intertemporal fundamental theorem of asset pricing
[51c]Step 1c: Non-linear pricing theory
[51c.1]Fundamental axioms
[51c.1.1]Law of one price
[51c.1.2]Non-linearity
[51c.1.3]Arbitrage
[51c.2]Valuation as evaluation
[51c.2.1]Variance and other shift principles
[51c.2.2]Certainty-equivalent principle
[51c.2.3]Distortion principles
[51c.2.4]Esscher principle
[51c.3]Intertemporal consistency
[51c.3.1]Continuous time variables
[51c.3.2]Non-linear “martingales”?
[51c.4]Point of interest and pitfalls
[51c.4.1]Linear (mis)uses of non-linear pricing
[51d]Step 1d: Valuation implementation
[51d.1]Equities
[51d.1.1]Discounted cash-flows
[51d.1.2]Multiples
[51d.2]Options
[51d.2.1]Bachelier
[51d.2.2]Black-Scholes
[51d.2.3]Heston
[51d.2.4]Valuation recipe
[51d.3]Fixed-income
[51d.3.1]Vasicek
[51d.3.2]Other models
[51d.3.3]Valuation recipe
[51d.4]Insurance
[51d.4.1]Life insurance
[51d.4.2]Non-life insurance
[51d.5]Real assets
[52]Step 2: Risk drivers identification
[52.1]Equities
[52.2]Fixed-income
[52.2.1]Rolling value
[52.2.2]Yield to maturity
[52.2.3]Alternative representations
[52.2.4]Parsimonious representations
[52.2.5]Spreads
[52.3]Derivatives
[52.3.1]Rolling value
[52.3.2]Implied volatility
[52.3.3]Alternative representations
[52.3.4]Parsimonious representations
[52.3.5]Risk drivers for a variance swap
[52.4]Commodities
[52.5]Credit
[52.5.1]Modelling default
[52.5.2]Ratings as risk drivers
[52.5.3]Risk drivers from conditioning
[52.6]Currencies
[52.7]Insurance
[52.8]Operations
[52.9]High frequency
[52.10]Strategies
[52.11]Points of interest, pitfalls, practical tips
[52.11.1]Spurious heteroscedasticity
[53]Step 3: Quest for invariance
[53a]Step 3a: Univariate quest for invariance
[53a.1]Efficiency
[53a.1.1]Heavy tails increments
[53a.1.2]Skewed and positive distributions
[53a.1.3]Stochastic volatility increments
[53a.1.4]Discrete increments
[53a.2]Trends
[53a.2.1]Deterministic trend
[53a.2.2]Stochastic trend
[53a.3]Seasonality
[53a.4]Short memory
[53a.5]Long memory
[53a.6]Volatility clustering
[53a.6.1]Price clustering
[53a.6.2]Time clustering
[53a.7]Discrete migrations
[53a.7.1]Markov chains
[53a.7.2]Structural models
[53a.8]Points of interest
[53a.8.1]Returns are not invariants
[53a.8.2]Sampling step size
[53b]Step 3b: Multivariate quest and forecasting
[53b.1]Mean-covariance approach
[53b.1.1]Mean reversion
[53b.1.2]Cointegration
[53b.1.3]Mean-covariance/analytical forecast
[53b.2]Probabilistic historical approach
[53b.2.1]Historical distribution
[53b.2.2]Historical forecast
[53b.3]Probabilistic copula-marginal approach
[53b.3.1]Static copula-marginal
[53b.3.2]Credit application
[53b.3.3]Dynamic copula-marginal
[53b.3.4]Copula-marginal forecast
[54b.4]Points of interest
[54b.4.1]Probabilistic, multivariate quest for invariance
[54b.4.2]Toward machine learning
[54b.4.3]Dynamic copula marginal forecast
[54b.4.4]Standardization
[54b.4.5]Non-synchronous data
[54b.4.6]Historical forecast with consecutive (non-)overlapping sequences
[54b.4.7]High-frequency volatility/correlation
[55]Step 4: Repricing
[55.1]Repricing functions
[55.1.1]Full repricing
[55.1.2]Carry
[55.1.3]Taylor approximation
[55.2]Techniques
[55.2.1]Scenario-based full repricing
[55.2.2]Analytical Taylor repricing
[55.2.3]Hybrid Taylor/full repricing
[55.2.4]Testing the repricing
[55.3]Equities
[55.3.1]Full repricing
[55.3.2]Carry
[55.3.3]Taylor approximation
[55.4]Fixed-income
[55.4.1]Zero-coupon bonds
[55.4.2]Coupon bonds
[55.4.3]Carry
[55.4.4]Taylor approximation
[55.5]Derivatives
[55.5.1]European call options
[55.5.2]Taylor approximation
[55.5.3]Variance swap
[55.5.4]Carry
[55.5.5]Taylor approximation
[55.6]Credit
[55.6.1]Full repricing
[55.6.2]Simplified regulatory framework
[55.7]Currencies
[55.7.1]Exchange rates
[55.7.2]Forward contracts
[55.7.3]Carry
[55.8]Pitfalls and practical tips
[55.8.1]Strategies
[55.8.2]Repricing and arbitrage
[55.8.3]Path dependence
[55.8.4]“Repricing”versus “asset pricing/valuation theory”
[55.8.5]Black-Scholes-Merton is exactly correct!
[55.8.6]Greeks for intra-day updates
[55.8.7]Greeks at the horizon
[55.8.8]Bond carry versus accrued interest
[55.8.9]Option carry versus theta
X. Risk management
[56]Step 5: Aggregation
[56a]Step 5a: Value aggregation
[56a.1]Portfolio value
[56a.1.1]Linear portfolio value
[56a.1.2]Sum-of-parts
[56a.1.3]Valuation recipe
[56a.1.4]Portfolio exposure
[56a.2]Portfolio weights
[56a.2.1]Generalized weights
[56a.2.2]Offset cash
[56a.3]Credit value adjustment
[56a.3.1]Counterparty credit risk exposure
[56a.3.2]Credit value adjustment computation
[56a.4]Liquidity value adjustment
[56a.5]Points of interest and pitfalls
[56a.5.1]Horizon-dependent exposure
[56a.5.2]Diffusive exposure
[56a.5.3]Solvency and collateral
[56b]Step 5b: Performance aggregation
[56b.1]Static market/credit risk
[56b.1.1]P&L
[56b.1.2]Returns
[56b.1.3]Benchmark
[56b.1.4]Scenario-probability distribution
[56b.1.5]Elliptical distribution
[56b.1.6]Quadratic-normal distribution
[56b.2]Dynamic market/credit risk
[56b.2.1]Portfolio rebalancing P&L
[56b.2.2]Allocation policy P&L
[56b.3]Stress-testing
[56b.3.1]Theory
[56b.3.2]Why have stress-tests
[56b.3.3]Panic copula
[56b.3.4]Extreme copula
[57]Step 6: Ex-ante evaluation
[57.1]Stochastic dominance
[57.2]Satisfaction/risk measures
[57.3]Mean-variance trade-off
[57.3.1]Mean
[57.3.2]Variance
[57.3.3]Standard deviation
[57.3.4]Mean-variance trade-off
[57.3.5]A strange success story
[57.4]The fundamental risk quadrangle
[57.4.1]Relevant cases
[57.4.2]Generalizations
[57.5]Expected utility and certainty-equivalent
[57.5.1]Common examples
[57.5.2]Computation
[57.6]Value at Risk and quantile
[57.6.1]Definition
[57.6.2]Computation
[57.7]Expected shortfall and sub-quantile
[57.7.1]Definition
[57.7.2]Computation
[57.8]Spectral/distortion satisfaction measures
[57.8.1]Definition
[57.8.2]Common examples
[57.8.3]Computation
[57.9]Coherent satisfaction measures
[57.9.1]Definition
[57.9.2]Common examples
[57.9.3]Computation
[57.10]Induced expectations
[57.10.1]Definition
[57.10.2]Common examples
[57.10.3]Computation
[57.11]Non-dimensional ratios
[57.11.1]Signal-to-noise ratio
[57.11.2]Downside ratios
[57.11.3]Correlation
[57.12]Pitfalls, points of interest and practical tips
[57.12.1]The Arrow-Pratt approximation of the certainty-equivalent
[57.12.2]Utility versus quantile
[57.12.3]Utility versus spectrum functions
[57.12.4]The Buhlmann and Esscher expectations are not distortion expectations
[57.12.5]Satisfaction measures under normality
[58]Step 7: Ex-ante attribution
[58a]Step 7a: Ex-ante performance attribution
[58a.1]Bottom-up exposures
[58a.1.1]Pricing factors
[58a.1.2]Style factors/smart beta
[58a.2]Top-down exposures: factors on demand
[58a.2.1]Analytical computation
[58a.2.2]Cardinality constraints
[58a.3]Relationship between bottom-up and top-down exposures
[58a.3.1]Subportfolios
[58a.4]Joint distribution
[58a.4.1]Elliptical distribution
[58a.4.2]Scenario-probability distribution
[58a.5]Application: hedging
[58a.6]Pitfalls and practical tips
[58a.6.1]Estimation versus attribution
[58a.6.2]The ex-ante attribution is not a regression on past data
[58b]Step 7b: Ex-ante risk attribution
[58b.1]General criteria
[58b.1.1]Isolated/“first in”proportional attribution
[58b.1.2]“Last in”proportional attribution
[58b.1.3]Sequential attribution
[58b.1.4]Shapley attribution
[58b.2]Euler decomposition
[58b.2.1]Standard deviation and variance
[58b.2.2]Certainty-equivalent
[58b.2.3]Quantile
[58b.2.4]Sub-quantile
[58b.2.5]Spectral satisfaction measures
[58b.2.6]Coherent measures
[58b.3]Linear attribution for induced expectations
[58b.3.1]Actuarial pricing
[58b.4]Minimum-torsion bets attribution of variance
[58b.4.1]Minimum-torsion bets
[58b.4.2]Effective number of bets
[59]Enterprise risk management
[59.1]General approach
[59.1.1]Portfolio: balance sheet
[59.1.2]Performance: income statement
[59.2]Banking regulatory framework
[59.2.1]Economic net income
[59.2.2]Default events
[59.2.3]Conditional losses
[59.2.4]Vasicek model
[59.2.5]Economic capital
[59.2.6]Risk attribution
[59.3]Insurance regulatory framework
[59.3.1]Economic net income
[59.3.2]Solvency capital requirement
[59.4]Points of interest
[59.4.1]CreditRisk+ approximation
XI. Portfolio management
[60]Step 8: Construction
[60a]Step 8a: Portfolio optimization
[60a.1]Mean-variance framework
[60a.1.1]Special portfolios
[60a.1.2]Quadratic target formulation
[60a.1.3]Linear target formulation
[60a.1.4]Setting the inputs
[60a.2]Analytical mean-variance
[60a.2.1]Total return
[60a.2.2]Excess return over risk-free
[60a.2.3]Excess return over benchmark
[60a.2.4]Total versus excess return
[60a.3]Numerical mean-variance
[60a.3.1]Constraints on positions/trade size
[60a.3.2]Constraints on number of positions
[60a.3.3]Transaction costs
[60a.4]Fundamental law of active management
[60a.4.1]Monetary impact of one signal
[60a.4.2]Information coefficient
[60a.4.3]Monetary impact of multiple signals
[60a.4.4]Aggregation
[60a.4.5]Transfer coefficient
[60a.5]Pitfalls, points of interest and practical tips
[60a.5.1]Black-Litterman equilibrium inputs via minimum relative entropy
[60b]Step 8b: Estimation and model risk
[60b.1]Mean-variance estimation risk measurement
[60b.1.1]Allocation as estimation
[60b.1.2]From predictive to estimative decisions
[60b.1.3]Two extreme allocation decisions
[60b.1.4]Decision theoretic allocation loss
[60b.2]Mean-variance Bayesian optimization
[60b.3]Mean-variance frequentist optimization
[60b.3.1]Tractable hypothesis set
[60b.3.2]Robust frontier
[60b.4]Probabilistic estimation risk measurement
[60b.4.1]Allocation as estimation
[60b.4.2]From predictive to estimative decisions
[60b.4.3]Probabilistic Bayesian optimization
[60b.4.4]Probabilistic frequentist optimization
[60b.4.5]Two-step approach
[60c]Step 8c: Cross-sectional strategies
[60c.1]Signals
[60c.1.1]Carry signals
[60c.1.2]Value signals
[60c.1.3]Technical signals
[60c.1.4]Fundamental and other signals
[60c.1.5]Signal processing
[60c.2]Premia
[60c.2.1]Signal-induced factor
[60c.2.2]Backtesting
[60c.3]Direct construction from signals
[60c.3.1]Signals as decision rules
[60c.4]Construction from signal predictions
[60c.4.1]Characteristic portfolio
[60c.4.2]Flexible factor
[60c.5]Relationship to APT
[60c.6]Multiple signals
[60c.6.1]Factor-mimicking portfolios
[60c.6.2]Relationship to APT
[60c.7]Points of interest, pitfalls, practical tips
[60c.7.1]Machine learning
[60d]Step 8d: Time series strategies
[60d.1]The market
[60d.1.1]Risky investment
[60d.1.2]Low-risk investment
[60d.1.3]Strategies
[60d.2]Expected utility maximization
[60d.2.1]The objective
[60d.2.2]Optimization
[60d.3]Option based portfolio insurance
[60d.3.1]Payoff design
[60d.3.2]Partial differential equation
[60d.3.3]Budget
[60d.3.4]Policy
[60d.3.5]A unified approach
[60d.4]Rolling horizon heuristics
[60d.4.1]Constant proportion portfolio insurance
[60d.4.2]Drawdown control
[60d.5]Signal induced strategy
[60d.6]Convexity analysis
[61]Step 9: Execution
[61.1]Market impact modeling
[61.1.1]Exogenous impact
[61.1.2]Endogenous impact
[61.2]Order scheduling
[61.2.1]Trading P&L decomposition
[61.2.2]Model P&L
[61.2.3]Moments of model P&L
[61.2.4]Model P&L optimization
[61.2.5]Quasi-optimal P&L distribution
[61.3]Order placement
[61.3.1]Step 1: Order scheduling
[61.3.2]Step 2: Order placement
[61.4]Microstructure signals
[61.4.1]Trade autocorrelation
[61.4.2]Order imbalance
[61.4.3]Price prediction
[61.4.4]Volume clustering
[61.5]Points of interest, pitfalls, practical tips
[61.5.1]Mean-variance optimization in complex models
[61.5.2]Price manipulation
[61.5.3]Testing
[62]Step 10: Ex-post performance analysis
XII. Finance toolbox
[63]Foundations
[63.1]Instrument value
[63.1.1]Fair value
[63.1.2]Transaction value
[63.1.3]Value versus price
[63.1.4]Exposure
[63.1.5]Leverage
[63.2]Portfolio value
[63.2.1]Long positions
[63.2.2]Short positions
[63.2.3]Generic positions
[63.3]Cashflows
[63.3.1]The jump rule
[63.3.2]Cumulative cashflows
[63.3.3]Re-invested cash-flows
[63.3.4]Cashflow adjusted value
[63.4]Market microstructure
[63.4.1]Limit order book
[63.4.2]Co-moving values
[63.4.3]Transaction variables
[63.4.4]Activity time
[63.4.5]Liquidity curve
[64]Performance definitions
[64.1]Profit-and-loss and payoff
[64.1.1]Profit-and-loss (P&L)
[64.1.2]Payoff
[64.2]Holding P&L of a position
[64.2.1]Long positions
[64.2.2]Short positions
[64.2.3]Generic positions
[64.3]Trading P&L of a position
[64.3.1]Single transaction
[64.3.2]Multiple transactions in one position
[64.4]Implementation shortfall
[64.5]Returns
[64.5.1]Basic definitions
[64.5.2]Generalized linear returns
[64.5.3]Excess returns
[64.5.4]Investments with capital injection
[64.5.5]Log-returns
[64.6]Path analysis
[64.7]Pitfalls and practical tips
[64.7.1]Linear versus compounded returns
[64.7.2]Multi-currency conversions
[64.7.3]Actual versus simple P&L
[65]Asset classes
[65.1]Equities
[65.2]Fixed-income
[65.2.1]Zero-coupon bond
[65.2.2]Bank account
[65.2.3]Coupon bond
[65.2.4]Interest rate swaps
[65.2.5]Amortizing financial instruments
[65.3]Derivatives
[65.3.1]Call/put option
[65.3.2]Futures
[65.3.3]Variance swaps
[65.4]Commodities
[65.5]Credit
[65.5.1]Default variables
[65.5.2]P&L in the presence of credit risk
[65.5.3]Spreads
[65.6]Foreign exchange
[65.6.1]Forward exchange rate
[65.6.2]Forward contracts
Case studies
XIII. Quantitative finance: the “Checklist”
[66]Monte Carlo Checklist
[66.1]Step 2: Risk drivers identification
[66.1.1]Market
[66.1.2]Credit
[66.2]Step 3: Quest for invariance
[66.2.1]Market
[66.2.2]Credit
[66.3]Step 4: Repricing
[66.4]Step 5: Aggregation
[66.5]Step 6: Ex-ante evaluation
[66.6]Step 7: Ex-ante attribution
[66.6.1]Ex-ante attribution: performance
[66.6.2]Ex-ante attribution: risk
[66.7]Step 8: Construction
[66.8]Step 9: Execution
XIV. Data science: factor models and learning
[67]Principal component analysis of the yield curve
[67.1]Cross-sectional structure of the yield curve covariance
[67.2]Finite set of times to maturity
[67.3]The continuum limit
[68]Machine learning for hedging
[68.1]Least squares regression
[68.1.1]Theoretical optimum
[68.1.2]Linear least squares regression
[68.1.3]Least squares regression tree
[68.2]Least absolute distance regression
[68.2.1]Theoretical optimum
[68.2.2]Linear least absolute distance regression
[68.2.3]Least absolute distance regression tree
[69]Machine learning for credit risk
[69.1]Credit default classification
[69.1.1]Background
[69.1.2]Fit and assessment
[69.1.3]Logistic regression
[69.1.4]Interactions
[69.1.5]Encoding
[69.1.6]Regularization
[69.1.7]Trees
[69.1.8]Gradient boosting
[69.1.9]Cross-validation
[70]Clustering for the stock market
[70.1]k-means clustering
[70.2]Shrinkage

24.6 Shrinkage PIC

Key points

  • Shrinkage estimators aim to reduce the variance of an estimator, and hence to improve its efficiency.
  • The most popular shrinkage estimator for the expectation is the James-Stein estimator, which averages the sample mean with a constant estimator (24.71).
  • Several covariance shrinkage estimators are available (Ledoit-Wolf (24.84), graphical lasso (24.102), factor analysis (24.105), etc.).

Shrinkage estimators average a high-variance estimate, say the sample mean, with a high-bias, subjective estimate of the value of the unknown parameter. The resulting shrinkage estimators trade off bias for efficiency, resulting in better estimators, according to the estimation assessment criteria that we discuss in Section 27.1.

In theory, we can apply shrinkage to the estimation of any property of a distribution. In practice, shrinkage is used mostly with expectations, covariances/correlations, and with factor loadings in linear factor models (14.14).

We cover the classical James-Stein shrinkage of the expectation parameters in Section 24.6.1.

In the next sections we present different techniques to construct shrinkage estimators of the covariance or correlation matrix: Ledoit-Wolf shrinkage ( Section 24.6.2); random matrix theory (Section 24.6.3); sparse eigenvectors rotations (Section 24.6.4); glasso over Markov networks (Section 24.6.5); factor analysis ( Section 24.6.6).

Another shrinkage methodology for the correlation matrix is the homogenous clusters correlation shrinkage, which we cover in the context of machine learning techniques, as an application of k-means clustering. Refer to Section 70.

We discuss the shrinkage of factor loadings in Section 25.5, fully dedicated to linear factor models.

Shrinkage is also a feature of Bayesian estimation, whereby data-driven estimates are naturally shrunk towards the Bayesian prior. We discuss Bayesian estimation for expectation and covariances in Section 24.5; and for factor loadings in Section 25.3

In Section 33.6 we discuss generalized shrinkage in the more general context of views processing.

24.6.1 Mean shrinkage: James-Stein

In this section we address the shrinkage estimation of the expectation (8.13) of a set of i.i.d. variables.

Historically, the first shrinkage estimator was the James-Stein estimator of the expectation, refer to the original article [Stein, 1955] and see also [Lehmann and Casella, 1998].

To understand the James-Stein shrinkage estimator, let us start from the sample mean ˆμ (24.116) of a time series of i.i.d. variables {ϵ1,…,ϵ¯t} (36.2).

The sample mean is unbiased but highly inefficient, see Example 27.10. In words, when we use as input different time series {ϵ1,…,ϵ¯t}, the output estimate varies widely, see Figure 23.15. As a result, “optimal” mean-variance allocations based on the sample mean vary wildly when different time series are used as input into the allocation process, see Section 60b.1.3.

Next, let us consider an estimator that is extremely efficient, namely a target vector μtarget which completely disregards any historical input data. A choice can be μtarget=0 as in [Alexander, 1998]. Alternatively, forward-looking targets implied by the notion of carry, see Section 55.1.2.

Then, as it turns out, we can improve the sample mean by “shrinking” it towards the shrinkage target: this way we obtain the James-Stein estimator

¯¯¯μ≡(1−γ)ˆμ+γμtarget,(24.71)

where γ∈(0,1) is a suitable constant that represents the confidence in our shrinkage target. The limit case γ=0 implies no shrinkage (full confidence in the sample mean), whereas the limit case γ=1 represents full confidence in the shrinkage target, as we disregard the sample mean.

For normally distributed variables εt∼N(μ,σ2) a sub-optimal confidence reads PIC 70.29 

γ=min{1, 2¯ttr(σ2)−2λ1(ˆμ−μtarget)'(ˆμ−μtarget)},(24.72)

where λ1 denotes the largest eigenvalue (43.412) of σ2. Notice how the confidence in the target decreases as the number of available observations increases: lim¯t→∞γ=0. Also note that the sub-optimal confidence (24.72) only applies when the dimension satisfies ¯ı≫2 [W], or else the shrinkage factor γ becomes negative. In practice, we can replace the unknown population parameters (σ2,λ1) by their sample counterparts (ˆσ2,ˆλ1) (24.3).

The shrinkage of the mean toward a prior target (24.71) also follows from a purely Bayesian approach to estimation, see Example 24.21.

PIC Example 24.21. In Example 24.13 on Bayesian theory, we consider a (univariate) normal i.i.d. variable εt (24.12), we compute the conditional likelihood f(i|μ,σ2) (24.40) and we assign a normal-inverse-Wishart model (24.45)-(24.47) to the prior distribution of the parameters (M,Σ2).
In Example 24.15 we then use the Bayesian approach to obtain the posterior distribution (23.172) of (M,Σ2), which turns out to be normal-inverse-Wishart itself (24.53)-(24.54). In this context, notice that the posterior location parameter μpos (24.56) mirrors the structure of the James-Stein shrinkage estimator (24.71), where we take as target the prior parameter μpri≡μtarget and we set the confidence as γ≡tpritpri+¯t: also note that, like in the optimal confidence expression (24.72), the confidence γ in the target μpri decreases as the number of observations ¯t increases (24.66).

PIC Example 24.22. Sample mean shrinkage: James-Stein estimator

PIC

Figure 24.14: Video

We consider the time series of log-returns ϵt of the stocks in the S&P 500 and we compute their sample mean ˆμ (24.116). We partition the S&P into the ¯k=10 MSCI US sector indices s1≡ “Energy”, s2≡ “Materials”, s3≡ “Industrials”, s4≡ “Consumer Discretionary”, s5≡ “Consumer Staple”, s6≡ “Healthcare”, s7≡ “Financials”, s8≡ “IT”, s9≡ “Telecommunications”, s10≡ “Utilities”.
As a simple example of target mean μtarget in (24.71) we can use the grand mean, which is the cross-sectional average of the entries of the sample mean

[μtarget]i=∑j∈skˆμj∑j∈sk1,      if i∈sk.(24.73)

Then we take γ=1 in formula (24.71) so that the shrinkage mean coincides with the target mean, ¯¯¯μ≡μtarget.
Figure 24.14 shows the sample mean ˆμ and the shrinkage mean ¯¯¯μ of the returns of two stocks in the same sector. Both means are computed over 100 different random small sub-samples of the whole sample. From the whole sample we compute the global mean, which represents the true “hidden” return, in the spirit of cross-validation, see Section 30.5. We observe that the sample mean varies widely across different sub-samples, i.e. it is highly inefficient. More formally, the average loss (27.47), or risk (27.46), for the sample mean is larger than for the shrinkage estimator. Figure 24.14 also shows the mean-variance optimal weights using shrinkage mean and sample mean. Notice that, using the sample mean, the allocation varies widely.

24.6.2 Covariance shrinkage: Ledoit-Wolf

The Ledoit-Wolf shrinkage in [Ledoit and Wolf, 2004] follows from the observation that the eigenvalues of the estimated covariance or correlation matrix tend to be more dispersed than the eigenvalues of the true data generating process.

To illustrate this effect, let us consider the case where the data generating process εt is standardized noise (36.2), i.e. all its entries are independent and identically distributed, with zero expectation and unitary variance

εt≡(ε1,t,…,ε¯ı,t)'∼fε i.i.d., with ⎧⎪ ⎪⎨⎪ ⎪⎩E{εi,t}=0V{εi,t}=1Cv{εi,t,εj,t}=0, for i≠j.(24.74)

We do not need to assume that the entries εi,t are normally distributed. However, we can consider normally distributed (49.1) noise as an example

εt≡(ε1,t,…,ε¯ı,t)'∼N(0¯ı×1,I¯ı) i.i.d..(24.75)

In this case, if we consider the spectral decomposition (43.372) of the true covariance

{λ,e}=eig(Cv{εt}),(24.76)

we obtain that all eigenvalues are equal to 1

Cv{εt}=I¯ı⇒λ=1¯ı.(24.77)

Suppose that we have a time series of realizations i={ϵ1,…,ϵ¯t} (23.1) of the standardized noise (24.74). Let us now consider the simplest estimator of the covariance, namely the sample covariance ˆσ2ε (24.117), and its spectral decomposition (43.372)

{ˆλ,ˆe}=eig(ˆσ2ε).(24.78)

Empirically, the sample spectrum ˆλ (24.78) is more dispersed, or steeper, than the true spectrum λ (24.76). To quantify this statement, let us define the condition number [W] of a dispersion matrix as the ratio of the largest to the smallest eigenvalue

κ(σ2)≡λ1λ¯ı.(24.79)

Then the sample condition number κ(ˆσ2) is larger than the true condition number

κ(ˆσ2)≡ˆλ1ˆλ¯ı>κ(σ2)≡λ1λ¯ı=1.(24.80)

In particular, let us consider the extreme, although in financial applications not rare, case where the number of observations ¯t is less than the dimension ¯ı of the variables εt. In this case, the sample covariance ˆσ2 becomes singular 70.24 

¯t<¯ı⇒rank(ˆσ2)=¯t<¯ı,(24.81)

which implies that at least the last eigenvalue is zero ˆλ¯ı=0. Hence the sample condition number goes to infinity, i.e. ˆσ2 is ill-conditioned

¯t<¯ı⇒κ(ˆσ2)=+∞.(24.82)

The degeneration of the condition number bears a geometrical interpretation, see Figure 24.15. The sample eigenvalues are more dispersed than the true eigenvalues, or equivalently the sample spectrum is steeper than the true spectrum. Hence, because the reference axes of a generic ellipsoid ∂E(μ,σ2) have length equal to the square root of the eigenvalues of σ2 (8.54), the ellipsoid ∂E(0,σ2) (12.42) of the true covariance σ2≡Cv{εt} is more spherical than the ellipsoid ∂E(0,ˆσ2) (12.42) of the sample covariance ˆσ2 (24.78), or intuitively

{λi}¯ıi=1⇔sphere   estimation→   {ˆλi}¯ıi=1⇔cigar.(24.83)

PIC Example 24.23. Estimated spectrum vs. true spectrum

PIC

Figure 24.15: Video

In Figure 24.15 we illustrate the average estimation error on the spectrum.

We consider a multivariate standard normal white noise εt (24.75) of large dimensions.

In the bottom right subplot we display the input, namely an ¯ıׯt panel i=(ϵ1|…|ϵ¯t) of joint i.i.d. realizations of the multivariate standard normal white noise (24.75). We draw one such panel i(j) multiple times for j=1,2,…,¯ȷ. For any given panel, we compute the sample covariance matrix ˆσ2(j) and extract the respective sample eigenvalues {ˆλ(j)i}i (24.78).

In the bottom left subplot we show the ¯ıׯı sample covariance matrix ˆσ2(j) (24.117) estimated from any given panel j.

In the top left plot we show the true eigenvalues {λ(j)i}i (24.77), sorted in decreasing order, as a function of the entry number i. Besides true spectrum, we also show the sample spectrum {ˆλ(j)i}i 24.78) sorted in decreasing order,and its average across different simulations j. The true spectrum is flat because all the entries are ones. However, the sample spectrum tilts, consistently with the Ledoit-Wolf effect.

In the top right plot we show the mean-covariance ellipsoid (8.47) of the first and last principal factors (8.82). The true principal factors are white noise, and therefore their mean-covariance ellipsoid is a circle (24.83). The mean-covariance ellipsoid of the sample principal factors is squeezed (24.83), because the largest principal axis is the square root of the largest sample eigenvalue, and similar for the smallest axis and sample eigenvalue, which are distorted by the Ledoit-Wolf effect.

As we change the input panel of realized white noise, the true covariance remains unaffected, but the sample covariance changes. However, as long as we do not change the number of joint realizations ¯t in the panel, the average sample spectrum and the average covariance ellipsoid remain unaffected. When we decrease the number of joint realizations ¯t, the average sample spectrum steepens and the average ellipsoid squeezes. In particular, when the number of joint realizations ¯t becomes less than the dimension of the white noise, we obtain the first zero sample eigenvalue: the sample covariance matrix is no longer invertible (24.81). Then, as we lower further the number of joint realizations ¯t, we obtain more zero eigenvalues, which is reflected in the right end of the average sample spectrum matching the horizontal axis, and the sample covariance ellipsoid flattening to a degenerate line.

The analysis is robust to distributional assumptions:  if we perform a similar analysis with exponentially distributed white noise, we observe the precise same phenomenon, both qualitatively and quantitatively, in both this sample spectrum and the sample covariance ellipsoid.

A shrinkage estimator that addresses the steepening of the spectrum (24.83) for an arbitrary covariance matrix is developed in [Ledoit and Wolf, 2004]. Similarly to the James-Stein shrinkage estimator of the mean (24.71), the shrinkage estimator of the covariance is a convex combination [W] of the sample covariance ˆσ2 and a target covariance, set as a multiple of the identity matrix σ2target≡1¯ıtr(ˆσ2)I¯ı

¯¯¯σ2=(1−γ)ˆσ2+γσ2target.(24.84)

In the case of normally distributed i.i.d. variables εt∼N(μ,σ2), the optimal coefficient γ∈[0,1] can be derived by minimizing the error ∥¯¯¯σ2−σ2∥2F, where ∥⋅∥F is the Frobenius norm (43.256), as explained in [Ledoit and Wolf, 2004] 70.30 . However it requires the knowledge of the true population covariance σ2, and therefore cannot be computed. [Ledoit and Wolf, 2004] propose an estimator for the theoretical optimal γ PIC

ˆγ=min{1,1¯t2∑¯tt=1tr{(ϵtϵ't−ˆσ2)2}tr{(ˆσ2−1¯ıtr(ˆσ2)I¯ı)2}},(24.85)

which asymptotically converges to the true coefficient. Notice how ˆγ decreases as the number of observations available in the sample mean increases: lim¯t→∞ˆγ=0.

The shrinkage estimator for the covariance (24.84) causes the spectrum to flatten. Hence, the shrinkage estimator is better conditioned [W] than the sample covariance 70.31 . The intuition behind the covariance shrinkage is as follows: the covariance shrinkage estimator (24.84) attempts to reverse the squeezing of the sphere caused by the estimation process (24.83).

24.6.3 Correlation shrinkage: random matrix theory

In Section 24.6.2 we used shrinkage to provide a more accurate spectral decomposition (43.413) of the true, unknown covariance matrix σ2≡Cv{εt} of a set of i.i.d. variables εt. More precisely, we flattened the spuriously steep spectrum of the sample estimate of the covariance matrix by averaging the sample estimate with a multiple of the identity matrix (24.84). A different approach to flatten spuriously steep spectrum of an estimate of the correlation matrix rests on random matrix theory, as discussed in Section 26.

Following [Rao, 2004], consider the Marchenko-Pastur distribution (26.300)

fMPΛ(x)≡(1−1q)δ(0)(x)1q>1+√(x−λ−)(λ+−x)2πxq1x∈(λ−,λ+),(24.86)

where we set

q≡¯ı¯t,λ±≡(1±√q)2,σ2≡1.(24.87)

Consider now a standardized noise εt that satisfies the assumptions (24.74). Given the vector ˆλ=(ˆλ1,…,ˆλ¯ı)' of sample eigenvalues, we can build its empirical pdf (26.36)

ˆfΛ(x|ˆλ)≡1¯ı∑¯ıi=1δ(ˆλi)(x).(24.88)

If the dimensions ¯ı and ¯t are large enough, then the empirical pdf (24.88) is approximated by the Marchenko-Pastur distribution (26.315)-(24.87)

ˆfΛ(x|ˆλ)¯ı≡q¯t→∞≈fMPΛ(x),(24.89)

see Example 24.24.

The approximation (24.89) holds regardless of the distribution of the white noise εt. Note that, when q≤1, the (closed) support [W] of the density fMPΛ (24.86) is the interval [λ−,λ+]; when q>1, the density fMPΛ (24.86) has the same support, plus a spike at zero. Therefore, according to the density fMPΛ, for any value of q the estimated sample eigenvalues ˆλ are scattered away from the true eigenvalues λ, which by construction are concentrated at 1.

It is possible to generalize the Marchenko-Pastur approximation (24.89) to the case of pure noise with general variance through its scaling property 70.22 

ˆfΛ(x|ˆλ)¯ı≡q¯t→∞≈1λfMPΛ(xλ),(24.90)

and also to the non-zero correlation case [Bun et al., 2017]. In the context of random matrix theory another notable limit distribution is Wigner’s semicircle distribution (26.210),as discussed in Section 26.5.

PIC Example 24.24. Correlation shrinkage: spectrum estimation via random matrix theory

PIC

Figure 24.16: Video

In Figure 24.16 we illustrate the power of the Marchenko-Pastur distribution fMPΛ (24.86).

Continuing from Example 24.23, in the top left plot we show the distribution of the eigenvalues, as approximated by a normalized histogram (49.281). The true distribution is peaked at 1 with no other eigenvalues. The sample distribution is dispersed around the peak of the true spectrum according to the Marchenko-Pastur distribution (24.89).

When we change the number of joint realizations ¯t, the Marchenko-Pastur distribution of the eigenvalues changes accordingly. In particular, when the number of joint realizations ¯t becomes less than the dimension of the white noise ¯ı, and we obtain more and more zero eigenvalues, the distribution of the eigenvalues peaks at zero.

This result is robust to distributional assumptions: when we perform a similar analysis with exponentially distributed white noise the Marchenko-Pastur distribution does not change.

Following [Bouchaud and Potters, 2009a] and [Bun et al., 2017], we can leverage random matrix theory for shrinkage purposes, as we proceed to explain. Consider a general data generating process εt≡(ε1,t,…,ε¯ı,t)' which generate data that are jointly i.i.d. across time and are a linear combinations of

- ¯k<¯ı signals {εsignlk}¯kk=1 independent of each other, with variances sorted in decreasing order (without loss of generality)

λ1≥⋯≥λ¯k;(24.91)

- ¯ı−¯k sources of noise {εsignlk}¯ık=¯k+1 independent of each and of the signals with smaller, equal variances λ¯k+1=⋯=λ¯ı=λ<λ¯k.

We can write the i.i.d. process εt as follows

⎛⎜⎝ε1⋅ε¯ı⎞⎟⎠=(e1,…,e¯k)×⎛⎜ ⎜⎝εsignl1⋅εsignl¯k⎞⎟ ⎟⎠+(e¯k+1,…,e¯ı)×o×⎛⎜ ⎜⎝εnoise¯k+1⋅εnoise¯ı⎞⎟ ⎟⎠,(24.92)

where

- e≡(e1,…,e¯ı) is an ¯ıׯı rotation/reflection (43.231) in R¯ı and thus it satisfies ee'=I¯ı;

- o an (¯ı−¯k)×(¯ı−¯k) rotation/reflection (43.231) in R¯ı−¯k and thus it satisfies oo'=I(¯ı−¯k).

The true spectrum (24.76) of the covariance matrix Cv{εt} of the i.i.d. process (24.92) coincides with the juxtaposition of the variances of both signals and noise, i.e.

eigval(Cv{εt})=(λ1,…,λ¯k,λ,…,λ)',(24.93)

while the eigenvectors (43.369) turn out to be the columns of matrix e 70.25 .

Now consider a time series for the i.i.d. process i={ϵ1,…,ϵ¯t}. For the signals, the sample eigenvectors and the sample eigenvalues are close to their true counterparts ˆλk≈λk and ˆek≈ek for k=1,…,¯k. Instead, the sample eigenvalues corresponding to the many sources of noise (k=¯ı−¯k+1,…,¯ı) will be affected by the elliptical distortion (24.83) illustrated in Figure 24.15.

Then the empirical pdf of the sample eigenvalues (24.88) splits into two components: the distribution of the large number of small eigenvalues (ˆλ¯k+1,…,ˆλ¯ı) stemming from the noise, which follows the rescaled Marchenko-Pastur law (24.90); and the spikes δ(ˆλk)(x) corresponding to the large eigenvalues (ˆλ1,…,ˆλ¯k) stemming from the signals

ˆfΛ(x|ˆλ)≈w×1¯k∑k≤¯kδ(ˆλk)(x)+(1−w)×1λnoisefMPΛ(xλnoise),(24.94)

where

- δ(λ)(x) denotes the Dirac delta (46.95) centered in λ, which spikes in x=λ;

- w=¯k∕¯ı is the relative weight of the two distributions: ¯k signal/spike eigenvalues, versus ¯ı−¯kMarchenko-Pastur/noise eigenvalues over a total of ¯ı eigenvalues;

- λnoise is the average of the noise eigenvalues λnoise≡1¯ı−¯k∑k>¯kˆλk; this number is the counterpart of λ in the true spectrum (24.93);

- fMPΛ(x) is the Marchenko-Pastur distribution (24.86), where q=(¯ı−¯k)∕¯t is the ratio of number of eigenvalues over observations, adapted from the ratio q≡¯ı¯t (24.87) to the current situation of ¯ı−¯k noise eigenvalues.

At this point we have all the theoretical instruments required to understand and perform spectrum shrinkage in practice. In particular, the decomposition (24.94) can be used to model the histogram of the sample spectrum ˆλ stemming from any input dataset of observations i={ϵ1,…,ϵ¯t}.

First of all, let us assume the observations ϵt to be already standardized, as explained in the joint time series (53b.56), in order to obtain correlation-like covariances that are immune to computational problems.

Then, we need to search for threshold ¯k such that the empirical distribution ˆfΛof the smaller eigenvalues ˆλk<ˆλ¯k best fits the rescaled Marchenko-Pastur law in the empirical pdf (24.94), at which point we cannot reject that the smaller sample eigenvalues stem from the noise component of process εt, and therefore they must be all equal

V:λ¯k+1=⋯=λ¯ı.(24.95)

The statistical interpretation of the view (24.95) is the following. The larger eigenvalues cannot have been generated by noise. However, for the smaller eigenvalues we cannot rule out that they stem from noise, i.e. a data generating process whose true covariance is a multiple λnoise of the identity matrix and thus has all equal eigenvalues.

The geometric interpretation of the view (24.95) is the following: the principal directions (8.71) ek identified by the smallest eigenvalues λk (k>¯k) of the covariance matrix σ2 are spurious and thus it makes more sense to assume isotropy, or spherical symmetry. Indeed, the decomposition is not affected by arbitrary rotations o applied to the noise component in the i.i.d. process (24.92).

To summarize, the spectrum shrinkage stemming from random matrix theory can be performed in four steps.

First, we perform the PCA decomposition (43.413) on the estimated sample covariance ˆσ2 (24.117), in order to compute the eigenvectors matrix ˆe (43.369) and the sample spectrum ˆλ (43.356).

Second, we determine the index ¯k such that the rescaled Marchenko-Pastur law fMPΛ (24.90) best fits the empirical distribution of the smaller eigenvalues ˆλk<ˆλ¯k: the fit can be measured by any divergence Dbetween the two distributions.

Third, we shrink the spectrum by imposing isotropy (24.95).

Fourth, we combine the eigenvectors ˆe with the spectrum ¯¯¯λ to obtain the new covariance matrix ¯¯¯σ2.

The corresponding routine reads:




Inputs ˆσ2



1. Perform spectral decomposition (ˆe,ˆλ)spectral dec.⇐ˆσ2=ˆeDiag(ˆλ)ˆe'(8.42)
for k=1,...,¯ı
2a. Compute empirical eigenvalue distribution ˆf(k)Λ⇐ˆλk+1,…,ˆλ¯ı
2b. Compute rescaled Marchenko-Pastur distribution˜f(k)Λ←1λnoisefMPΛ(xλnoise)
end
3. Assess divergence for optimal cutoff ¯k←argmink∈{1,...,¯ı−1}D(˜f(k)Λ∥ˆf(k)Λ)
4. Impose isotropy ⎧⎨⎩¯λi←ˆλi for all i≤¯k¯λi←1¯ı−¯k∑k>¯kˆλk for all i>¯k
5. Compute new covariance ¯¯¯σ2←ˆeDiag(¯¯¯λ)ˆe'



Outputs ¯¯¯σ2



Table 24.9: Spectrum shrinkage stemming from random matrix theory

The spectrum shrinkage routine (Table 24.9) can be applied at first impression to any covariance matrix ˆσ2. However, if the input is not the sample covariance ˆσ2 (24.117) of a time series of i.i.d variables {ϵ1,…,ϵ¯t}, the optimization in Step 3 must be modified accordingly, see e.g. [Potters et al., 2005].

PIC Example 24.25. Spectrum shrinkage via random matrix theory

PIC

Figure 24.17: Video

In Figure 24.17 we highlight the main features of the covariance shrinkage based on the Marchenko-Pastur distribution from random matrix theory.

In the top-left plot we show the sample spectrum (43.389) computed from the sample covariance (24.78). Then we assume that the smallest eigenvalues are noisy and thus their relative size bears no meaning and we shrunk their value to their mean (24.95). Here we show the empirical distribution of the small sample eigenvalues.

In the top-center plot we also show the rescaled Marchenko-Pastur distribution (24.90) that is supposed to fit the noisy sample eigenvalue distribution (24.94). Furthermore, we highlight the spike toward which the noisy eigenvalues are shrunk which is supposed to be their true population spectrum.

In the top-right plot we show the mean-covariance ellipsoids (8.47) defined by the sample covariance and by the shrinkage covariance respectively, sliced along the plane spanned by a high-volatility direction, namely the 10-th sample eigenvector; and a low-volatility direction, namely the last eigenvector. The direction of the last eigenvector is statistically the same as the direction of all the eigenvectors whose value has been shrunk, due to the isotropy implied by the shrinkage to the mean (24.95).

In the bottom plot we display all the original and filtered correlations sorted in decreasing order. The overall effect of the eigenvalue shrinkage is a mild shrinkage of the correlations toward zero.

As we vary the cutoff that divides noise from the signal, the fit of the Marchenko-Pastur distribution to the distribution of the noisy empirical eigenvalues varies. We need to select a value that can be accepted as consistent Marchenko-Pastur distribution.

When all the eigenvalues are shrunk to isotropy, we obtain the Ledoit-Wolf shrinkage target (24.84).

24.6.4 Covariance shrinkage: sparse eigenvector rotations

In Sections 24.6.2-24.6.3 we used shrinkage to provide an accurate spectral decomposition (43.413) of the true, unknown covariance matrix of a set of i.i.d. variables εt (36.2)

Cv{εt}≡σ2=eDiag(λ)e'.(24.96)

More precisely, we flattened the spuriously steep spectrum of the sample estimate of the covariance matrix by averaging the sample estimate with a suitable target matrix (24.84), or by averaging the smallest eigenvalues (24.95).

The approach in [Cao et al., 2011] to improve the estimate of the spectral decomposition (24.96) acts on the eigenvectors matrix e, rather than on the spectrum of eigenvalues λ (43.356).

We start with the sample covariance ˆσ2≡ˆeDiag(ˆλ)ˆe' (24.117) of a set of i.i.d. variables εt≡(ε1,t,…,ε¯ı,t)', computed from a time series i={ϵ1,…,ϵ¯t} of independent realizations of εt. We stress the fact that it is always advisable to standardize (54b.83) the observations ϵt before focusing on the estimation of the covariance, in order to obtain correlation-like estimates that are immune to computational problems.

To define the shrinkage views V, let us recall that the eigenvectors matrix e is orthogonal (43.398) and hence represents a rotation (or a reflection) [W]. Then we set shrinkage views imposing that the eigenvectors matrix e is a sequence of ¯k simple rotations along the coordinate planes

V:e=r(i1,i'1,α1)×⋯×r(i¯k,i'¯k,α¯k),(24.97)

where each rotation is a Givens rotation [W], that only involves two indices and reads entrywise

r(i,j,α)≡⎧⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪⎨⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪⎩cosαentries i, i and j, jsinαentry i, j−sinαentry j, i1all remaining diagonal elements0all remaining elements..(24.98)

We then define the shrinkage estimate of the covariance as the matrix that maximizes the conditional likelihood f(i|σ2) (24.38) of the time series i under the normal distribution assumption εt∼N(μ,σ2), all while satisfying the constraint (24.97) on the eigenvectors

¯¯¯σ2=argmaxσ2f(i|σ2)(24.99)subject to {σ2≡eDiag(λ)e'e∈V.

Maximizing the target in (24.99) is equivalent to finding the eigenvectors e that minimize the following “distance” with respect to the sample eigenvectors ˆe 70.26 

D(e,ˆe)≡det(Diag(diag(e'ˆσ2e))).(24.100)

Then the eigenvectors that determine shrinkage estimate σ2, which satisfies the shrinkage views V (24.97) and at the same time is as close as possible to the original estimate ˆσ2, become

¯e≡argmine∈VD(e,ˆe).(24.101)

The shrinkage (24.101) under the rotation views (24.97) is not computationally tractable. However, [Cao et al., 2011] provide a heuristic approximation to the true optimum.

24.6.5 Covariance shrinkage: glasso

In this section we present an approach that shrinks the covariance Cv{ε} (8.16) and correlation Cr{ε} (10.15) matrices of a set of i.i.d. variables (36.2) to a parsimonious undirected graphical model or Markov random field (21.82).

As explained in Section 21.7, a set of simultaneous i.i.d. variables εt=(ε1,…,ε¯ı) is a parsimonious Markov random field when the number of 0-entries in the inverse correlation ϕ2≡(c2)−1 is large.

PIC Example 24.26. Refer to Example 21.5 to visualize a simple graphical model.

We start from the sample inverse correlation with flexible probabilities ˆϕ2HFP, i.e. the inverse of the correlation matrix ˆc2HFP (24.8), estimated on a time series i={ϵ1,…,ϵ¯t} of realizations of the i.i.d. variables εt.

In order to shrink the market to a parsimonious network, we impose the view that the inverse correlations generate a parsimonious network with few connections. Hence, we want the null entries of the inverse correlations to be at least a given lower boundary k

V≡{ϕ2:(num of 0-entries in ϕ2)≥k}.(24.102)

We then find the parameter ¯¯¯ϕ2 that minimizes the distance from ˆϕ2HFP, all while satisfying the view (24.102)

¯¯¯ϕ2≡argminϕ2∈VD(ϕ2,ˆϕ2HFP).(24.103)

where D is the pseudo-distance (E.22.208). The target in (24.103) is equivalent to the relative entropy under normal assumptions (5.75).

The shrinkage views (24.102) are a cardinality constraint (45.138). Minimizing the distance (E.22.208) under such a constraint is impossible in practice. However, we can resort to quasi-optimal heuristics. One such heuristic is via greedy algorithms [W], in particular forward selection, see Section 47.2.3.

A moderately less optimal, although more efficient, solution is the graphical lasso or glasso, see [Hastie et al., 2009]. Similarly to the regular lasso discussed in Section 25.5, the graphical lasso minimizes the distance (E.22.208) from a generic symmetric and positive definite matrix c––2, while penalizing the sum of the absolute values of the entries ∥ϕ2∥1≡∑i,j|[ϕ2]i,j|, as follows

ϕ2λ≡argminϕ2tr(ϕ2c––2)−lndet(ϕ2)+λ∥ϕ2∥1.(24.104)

As in the standard lasso, the penalty λ forces some entries of ϕ2 to shrink to zero, and the number of zero entries increases as the penalty λ increases, see Example 24.27.

PIC Example 24.27. Stocks Markov network structure via graphical lasso of correlation

PIC

Figure 24.18: Video

We consider the same setting as in Example 24.19. We apply the glasso shrinkage, for different penalty values λ, on the inverse ˆϕ2 of the sample correlation ˆc2 stemming from ˆσ2 (10.22), as described in Table 24.10.
In the top plot of Figure 24.18 we show the evolution of the graph of the shrunk inverse correlation matrix ϕ2λ (24.104).
In the bottom-left plots we display the heat map of the HFP correlation matrix ˆs2←corr(ˆs2) (24.3)-(10.22) and its inverse ˆϕ2.
In the bottom-right plots we display the heat map of the correlation matrix stemming from the regularized inverse covariance ϕ2λ (24.104).
The elements of the inverse-correlation terms become null one at a time as the penalty increases. When all the elements of a column/row (except for the diagonal element) of the inverse matrix have become null, then the same column/row disappears also in the covariance matrix. This is a consequence of the formula on partitioned matrix inversion (43.615).
To implement the glasso we used the code adapted by H. Karshenas from the original code provided by [Friedman et al., 2008].
Refer to Example 24.19 to see how the same parsimonious structure of the Gaussian Markov random field (21.85) can be obtained through Bayesian shrinkage.

To summarize, the routine to perform shrinkage to a parsimonious network proceeds as follows




Inputs σ–––2,k,{λι}



1. Compute correlation c––2←Diag(1¯ı×1.∕σ–––vol)σ–––2Diag(1¯ı×1.∕σ–––vol)
for ι=1,...,¯ı
2. Perform graphical lasso˜ϕ2λιgraphical lasso⇐(c––2,λι) (24.104)
3. Extract correlation ⎧⎪ ⎪ ⎪ ⎪⎨⎪ ⎪ ⎪ ⎪⎩˜c2λι←(˜ϕ2λι)−1c2λι←Diag(1¯ı×1.∕[˜cλι]vol) ˜c2λι Diag(1¯ı×1.∕[˜cλι]vol)ϕ2λι←(c2λι)−1=Diag([˜cλι]vol) ˜ϕ2λι Diag([˜cλι]vol)
end
4. Compute optimal penalty¯λ←min{λι such that (num of 0-entries in ϕ2λι)≥k}
5. Shrunk inv. correlation ¯¯¯ϕ2←ϕ2¯λ
6. Shrunk correlation ¯c2←(¯¯¯ϕ2)−1
7. Shrunk covariance ¯¯¯σ2←Diag(σ–––vol)¯c2Diag(σ–––vol)



Outputs ¯¯¯σ2,¯c2,¯¯¯ϕ2



Table 24.10: Markov network shrinkage

24.6.6 Covariance shrinkage: factor analysis

In this section we present factor analysis, used to shrink the covariance matrix Cv{ε} (8.16) or correlation matrix Cr{ε} (10.15) of a set of i.i.d. variables (36.2) to a low-rank-diagonal structure (43.608).

More precisely, let us define the low-rank diagonal function

lrd(b,d)≡bb'low-rank+Diag(d)diagonal,(24.105)

where b is a full-rank ¯nׯk matrix (typically ¯k≪¯n) and d is an ¯n×1 positive vector.

Then, given a reference ¯nׯn symmetric, positive-definite matrix σ2, e.g. the HFP covariance (24.3), the output ¯¯¯σ2 is a low-rank diagonal matrixlrd(β,δ) (24.105) determined by the matrix β and vector δ that solve the following optimization problem

(β,δ)≡argminb,dD(lrd(b,d),σ2).(24.106)

for a suitable divergence D.

If the divergence D is specified as the distance induced (43.263) by the Frobenius norm (43.256), the optimization is solved via the principal axis factorization, as discussed in Section 14.4.2.

If D is specified as the relative entropy between normal variables (5.75), the optimization is solved by the maximum likelihood factorization, as discussed in Section 21.3.2.

Refer also to [MacCallum et al., 2007] and [Cudeck and MacCallum, 2012] for a discussion on the performance of the two algorithms.

PIC Example 24.28. Let us consider the following 3×3 positive-definite covariance matrix

σ2≡⎛⎜⎝3.100.360.160.364.310.280.160.280.88⎞⎟⎠.(24.107)

We shrink σ2 towards a low-rank diagonal matrix with ¯k=1 (24.105) using the principal axis factorization algorithm (PAF) (Table 14.3)

σ2;PAF≡lrd(βPAF,δPAF)≈σ2;(24.108)

and the maximum likelihood factorization algorithm (MLF) (Table 21.1)

σ2;MLF≡lrd(βMLF,δMLF)≈σ2.(24.109)

Then we compare the distance of each factor analysis matrix σ2;PAF and σ2;MLF with respect to the original covariance matrix σ2 using both i) the Frobenius distance (43.263)-(43.256) and ii) the (normal) relative entropy (5.75)

Frobenius distanceRelative Entropy––––––––––––––––––––––––––––––––––––––––––––––PAF10−6×0.1810−14×0.13MLF10−6×0.1910−15×0.9(24.110)

PIC

¯¯¯μ≡(1−γ)ˆμ+γμtarget,
¯¯¯σ2=(1−γ)ˆσ2+γσ2target.
V≡{ϕ2:(num of 0-entries in ϕ2)≥k}.
lrd(b,d)≡bb'low-rank+Diag(d)diagonal,
⎛⎜ ⎜ ⎜ ⎜ ⎜ ⎜⎝X1⋅Xn⋅X¯n⎞⎟ ⎟ ⎟ ⎟ ⎟ ⎟⎠X=⎛⎜ ⎜ ⎜ ⎜ ⎜ ⎜⎝α1⋅αn⋅α¯n⎞⎟ ⎟ ⎟ ⎟ ⎟ ⎟⎠α+⎛⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜⎝β1,1⋯β1,k⋯β1,¯k⋅⋅⋅βn,1⋯βn,k⋯βn,¯k⋅⋅⋅β¯n,1⋯β¯n,k⋯β¯n,¯k⎞⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟⎠β⎛⎜ ⎜ ⎜ ⎜ ⎜ ⎜⎝Z1⋅Zk⋅Z¯k⎞⎟ ⎟ ⎟ ⎟ ⎟ ⎟⎠Z+⎛⎜ ⎜ ⎜ ⎜ ⎜ ⎜⎝˚ε1⋅˚εn⋅˚ε¯n⎞⎟ ⎟ ⎟ ⎟ ⎟ ⎟⎠˚ε,
E{X}≡⎛⎜ ⎜ ⎜ ⎜ ⎜ ⎜⎝E{X1}⋅E{Xn}⋅E{X¯n}⎞⎟ ⎟ ⎟ ⎟ ⎟ ⎟⎠,
ˆμ≡esˆμ(i)≡1¯t∑¯tt=1ϵt;
εt∼F(ε) i.i.d..
λ1≥⋯≥λ¯n≥0.
γ=min{1, 2¯ttr(σ2)−2λ1(ˆμ−μtarget)'(ˆμ−μtarget)},
ˆs2HFPε≡ˆCvHFP{ε}=∑¯tt=1pt(ϵt−ˆmHFPε)(ϵt−ˆmHFPε)'(24.3)=∑¯tt=1ptϵtϵ't−ˆmHFPε(ˆmHFPε)'.
εt∼N(μ,σ2).
f(i|θ)=exp(¯t(θ'μˆημ+tr(θσˆησ)−ψ(θ)+lnh))(24.40)=(2π)−¯ı¯t2det(σ2)−¯t2exp(−¯t2[tr(ˆs2(σ2)−1)+(ˆm−μ)'(σ2)−1(ˆm−μ)]),
(Σ2)−1∼Wishart(νpri,1νpri(σ2pri)−1);
M|σ2∼N(μpri,1tpriσ2),
fpos(θ)≡f(θ|i)=f(i|θ)fpri(θ)∫f(i|ϑ)fpri(ϑ)dϑ,
(Σ2)−1|i∼Wishart(νpos,1νpos(σ2pos)−1),
M|σ2,i∼N(μpos,1tposσ2).
μpos=γμμpri+(1−γμ)ˆm,
ˆmlarge dataset ¯t←μcl_eq=μposhigh confidence tpri→μpri.
ℓoss(θ,ˆq)≡∥q(θ)−ˆq∥22.
r(θ,es)≡E{D(q(Θ)∥es(I))|θ}.
X∼N(μ,σ2),
eig(q)≡{eigval(q),eigvec(q)}={λ,e},
i≡(ϵ1,…,ϵt,…,ϵ¯t)≡⎛⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜⎝ϵ1,1ϵ1,tϵ1,¯t⋅⋅⋅ϵi,1⋯ϵi,t⋯ϵi,¯t⋅⋅⋅ϵ¯ı,1ϵ¯ı,tϵ¯ı,¯t⎞⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟⎠,
εt≡(ε1,t,…,ε¯ı,t)'∼fε i.i.d., with ⎧⎪ ⎪⎨⎪ ⎪⎩E{εi,t}=0V{εi,t}=1Cv{εi,t,εj,t}=0, for i≠j.
ˆσ2≡esˆσ2(i)≡1¯t∑¯tt=1(ϵt−ˆμ)(ϵt−ˆμ)'.
{ˆλ,ˆe}=eig(ˆσ2ε).
{λ,e}=eig(Cv{εt}),
√eigval(Cv{X})=principal semi-axes length⇔shape of ellipsoid.
∂E(Mod{X},MoDis2{X})≡{x:(x−Mod{X})'(MoDis2{X})−1(x−Mod{X})=1}.
εt≡(ε1,t,…,ε¯ı,t)'∼N(0¯ı×1,I¯ı) i.i.d..
Cv{εt}=I¯ı⇒λ=1¯ı.
∂E(E{X},Cv{X})={x∈R¯n:(x−E{X})'(Cv{X})−1(x−E{X})=1}.
ZPCn=argmaxZ∈C{0,...,n−1}V{Z},
{λi}¯ıi=1⇔sphere   estimation→   {ˆλi}¯ıi=1⇔cigar.
¯t<¯ı⇒rank(ˆσ2)=¯t<¯ı,
∥a∥F≡∥vec(a)∥2=√tr(a'a)=√∑min(¯n,¯¯¯m)n=1γ2n,
s2=e×Diag(λ)×e',
¯fWOEX(λ)=fMPΛ(λ)≡(1−1q)δ(λ)1q>1+√(λ−λ−)(λ+−λ)2πλqσ21λ∈(λ−,λ+),
ˆfΛ(y|λ)≡1¯n∑¯nn=1δ(y−λn),
ˆfΛ(x|ˆλ)≡1¯ı∑¯ıi=1δ(ˆλi)(x).
lim¯n→∞ˆf(λ|λ(xMP))=fMPΛ(λ).
q≡¯ı¯t,λ±≡(1±√q)2,σ2≡1.
ˆfΛ(x|ˆλ)¯ı≡q¯t→∞≈fMPΛ(x),
fMPΛ(x)≡(1−1q)δ(0)(x)1q>1+√(x−λ−)(λ+−x)2πxq1x∈(λ−,λ+),
¯fGOEX(λ)=fSCΛ(λ)≡1πσ2√2σ2−λ21λ∈(−√2σ,√2σ).
ˆhistX(x)≡∑jˆf(j)Δ1x∈X(j).
o∈O(¯n)⇔o'o=oo'=I¯n.
⎛⎜⎝ε1⋅ε¯ı⎞⎟⎠=(e1,…,e¯k)×⎛⎜ ⎜⎝εsignl1⋅εsignl¯k⎞⎟ ⎟⎠+(e¯k+1,…,e¯ı)×o×⎛⎜ ⎜⎝εnoise¯k+1⋅εnoise¯ı⎞⎟ ⎟⎠,
e≡⎛⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜⎝e1,1⋮en,1⋮e¯n,1∣∣ ∣ ∣ ∣ ∣ ∣ ∣∣⋯∣∣ ∣ ∣ ∣ ∣ ∣ ∣∣e1,n⋮en,n⋮e¯n,n∣∣ ∣ ∣ ∣ ∣ ∣ ∣∣⋯∣∣ ∣ ∣ ∣ ∣ ∣ ∣∣e1,¯n⋮en,¯n⋮e¯n,¯n⎞⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟⎠≡(e1|⋯|en|⋯|e¯n),
ˆfΛ(x|ˆλ)¯ı≡q¯t→∞≈1λfMPΛ(xλ),
∫Tδ(u)(t)g(t)dμ(t)=g(u).
eigval(Cv{εt})=(λ1,…,λ¯k,λ,…,λ)',
ˆfΛ(x|ˆλ)≈w×1¯k∑k≤¯kδ(ˆλk)(x)+(1−w)×1λnoisefMPΛ(xλnoise),
˜ϵi,t≡Φ−1ν(ui,t),
V:λ¯k+1=⋯=λ¯ı.
en≡argmaxv∈V{0,...,n−1}V{v'X},
λ≡⎛⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜⎝λ1⋮λ¯kλ¯k+1⋮λ¯n⎞⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟⎠≡⎛⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜⎝λ1⋮λ¯kλRe1+iλIm1λRe1−iλIm1⋮λRe¯ȷ+iλIm¯ȷλRe¯ȷ−iλIm¯ȷ⎞⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟⎠.
Cv{X}=e×Diag(λ)×e',
λn∈R,
Cv{εt}≡σ2=eDiag(λ)e'.
˜εi,t≡εi,t−ˆμiˆσi,
ee'=e'e=I¯n.
f(i|θ)i.i.d.=∏¯tt=1f(ϵt|θ),
V:e=r(i1,i'1,α1)×⋯×r(i¯k,i'¯k,α¯k),
¯¯¯σ2=argmaxσ2f(i|σ2)(24.99)subject to {σ2≡eDiag(λ)e'e∈V.
¯e≡argmine∈VD(e,ˆe).
Cv{X}≡⎛⎜ ⎜ ⎜ ⎜⎝V{X1}Cv{X1,X2}⋅Cv{X1,X¯n}Cv{X2,X1}V{X2}⋅⋅⋅⋅⋅⋅Cv{X¯n,X1}⋅⋅V{X¯n}⎞⎟ ⎟ ⎟ ⎟⎠.
c2X≡Cr{X}≡⎛⎜ ⎜ ⎜ ⎜⎝1Cr{X1,X2}⋅Cr{X1,X¯n}Cr{X2,X1}1⋅Cr{X2,X¯n}⋅⋅⋅⋅Cr{X¯n,X1}⋅⋅1⎞⎟ ⎟ ⎟ ⎟⎠.
(Ym⊥⊥Yn)|Y−{m,n}⇔(m∼n)∉E,
ˆc2HFPε=corr(ˆs2HFPε)(10.22)=Diag(1¯ı×1.∕ˆsHFPε;stdev) ˆs2HFPε Diag(1¯ı×1.∕ˆsHFPε;stdev).
D(ϕ2,ϕ–––2X)≡tr(ϕ2(ϕ–––2X)−1)−ln(det(ϕ2)).
¯¯¯ϕ2≡argminϕ2∈VD(ϕ2,ˆϕ2HFP).
DKL(μ1,σ21∥μ0,σ20)=12(tr(σ21(σ20)−1)−ln(det(σ21(σ20)−1))+(μ1−μ0)'(σ20)−1(μ1−μ0)−¯n).
k––≤card({n such that xn≠0})≤¯k,
corr(σ2)≡Diag(1¯n×1vol(σ2))×σ2×Diag(1¯n×1vol(σ2)),
ϕ2λ≡argminϕ2tr(ϕ2c––2)−lndet(ϕ2)+λ∥ϕ2∥1.
[a−1]≠n,n=−[a−1]n,n[a]−1≠n,≠n[a]≠n,n.
Y∼N(μ,σ2),
s2¯n=βs2¯kβ'+Diag(δ⊙δ),
D(x,y)≡∥x−y∥,

Bibliography


External Links


Discussions


 
arpm small logo
About us Start here
Clients and partners Corporate program Academia program
Contact us Events Book
Contact us   Linkedin
Terms ⚪ Privacy policy ⚪ Refund policy ⚪ Cookies policy ⚪ Copyright ⚪ IT requirements
© 2026 ARPM, All rights reserved.