24.2 Maximum likelihood


Key points
- The maximum likelihood with flexible probabilities estimators of the mean and covariance under normality (24.12) coincide with the historical with flexible probabilities estimators (24.18).
- The maximum likelihood with flexible probabilities estimators of the location and dispersion parameters of a Student distribution (24.19) with given degrees of freedom can be computed by means of the recursive routine in Table 24.1.
In this section we discuss how to estimate the location and dispersion of the distribution of i.i.d. variables by applying the maximum likelihood approach introduced in Section 23.4. We consider two parametric assumptions: normal (Section 24.2.1) and Student (Section 24.2.2).
24.2.1 Normal assumption
In this section we apply the maximum likelihood approach for exponential family distributions introduced in Section 23.4.5 to the special case of the multivariate normal (49.1)
| | (24.12) |
The standard parameters are (-dimensional expected value) and ( positive definite and symmetric covariance matrix), which are the most important properties of location and dispersion (Section 8).
The multivariate normal is an exponential family distribution with natural parameters (49.413), features (49.417) and log-partition function (49.421). The natural parameters are in one-to-one correspondence with the standard parameters . Our goal is to estimate and therefore .
Given a realized time series consider the historical mean with flexible probabilities (24.1)
| | (24.13) |
and the historical covariance matrix with flexible probabilities (24.3)
| | (24.14) |
Then the historical with flexible probabilities (HFP) sample features (23.80) read as follows
|
| (24.15) |
From the one-to-one correspondence (49.413), the log-likelihood with flexible probabilities (23.79) reads 70.2
We could proceed to maximize the likelihood, as in (23.67). However, it is easier to leverage the general results for exponential family variables from Section 23.4.5. The gradient of the log-partition is (5.68). Then the maximum likelihood estimates of the natural parameters (23.82) follow from inverting the gradient 70.37
| | (24.17) |
By expressing the standard parameters in terms of the canonical parameters (49.417), we obtain that the maximum likelihood estimates of the location and dispersion for normal variables are the historical mean and historical covariance with flexible probabilities (24.13)-(24.14)
|
| (24.18) |
which also follows from the limit of maximum likelihood under Student assumptions (24.24).
It is comforting to see that we obtain the “same” estimates using two completely different principles, namely the historical and maximum likelihood. However, keep in mind that the assumptions of the maximum likelihood approach, being parametric (23.59), are more restrictive.
24.2.2 Student
assumption
In this section we apply the maximum likelihood approach to estimate location and dispersion under the Student distribution assumption.
Let us assume as parametric specification (23.59) for the i.i.d. variables a Student distribution
| | (24.19) |
where the pdf reads (49.144)
| | (24.20) |
In this context is the space of parameters , where is a -dimensional location vector; is the positive definite and symmetric dispersion matrix; and is a positive scalar representing the degrees of freedom, which determine the thickness of the tails. Notice the power law decay in the tails, determined by the degrees of freedom .
When the power law becomes an exponential decay, and the Student becomes the normal distribution (49.1). As decreases, the tails become thicker and the probability of extreme scenarios increases. When the Student becomes the Cauchy distribution (49.164), whose tails are so thick that neither expected values nor covariances are defined, and yet the location and dispersion parameters are always well defined.
Example 24.2. Location-dispersion: MLFP ellipsoid
Video 
Let us consider the same setting as in Example 24.1.
In Figure 24.2 we show the realized residuals
of a GARCH
(41.75)-(41.77) fit of two stocks’ returns. We overlay the ellipsoid
(8.45) defined by the MLFP
estimates of location
(24.21) and dispersion
(24.22), computed with the routine in Table 24.1, with
.
We also plot the respective flexible probabilities
and display the effective
number of scenarios
(23.21).
The Student distribution does not belong to the exponential family (49.395), unless (24.12). Hence, we cannot leverage the general results in Section 23.4.5. However, we can explicitly perform the maximum likelihood with flexible probabilities (MLFP) optimization (23.67), obtaining the MLFP parameter as the solution of the following implicit equation 70.4
| | (24.21) |
and the MLFP parameter as the solution of
| | (24.22) |
with weights defined as
| | (24.23) |
Note that the scatter matrix estimator (24.22) is purposely not normalized by because is not the covariance (49.154). On the other hand, the location vector estimator (24.21) is normalized, because is the expectation (49.153), when the expectation is properly defined ().
In theory, to calibrate the degrees of freedom , we should compute a solution for each value of , then compute the respective likelihood, and finally choose the degrees of freedom that give rise to the highest likelihood (23.67). In reality, we recommend using a small number of degrees of freedom, such as , regardless of the fit, for robustness (for a detailed discussion on robustness, see Section 23.7).
Similarly, it is possible to perform the MLFP optimization under the most general multivariate elliptical assumption, which include the Student (24.19) as a special case 70.3 .
The MLFP parameters and depend on the assumption we make on the degrees of freedom.
In the normal case () the weights (24.23) become constant, , and assume the same expression than the historical with flexible probabilities (HFP) mean (24.1) and covariance (24.3)
| | (24.24) |
This follows also from a direct maximization of the normal likelihood (24.18).
In the general case of finite degrees of freedom , the weights (24.23) depend on and via the Mahalanobis distance (43.185). Hence, extreme scenarios, i.e. observations with large Mahalanobis distance, , have a negligible weight, , in the estimation. This result is intuitive: with low degrees of freedom , the tails are thick and thus we expect to see several extreme scenarios. If we did not down-weight such scenarios, i.e. if we set , the MLFP parameters would be the same as the HFP estimates and thus the ensuing ellipsoid would be spuriously stretched out to account for such scenarios, as we saw in the case of the geometrical interpretation (24.11).
The MLFP formulas (24.21)-(24.22)-(24.23) are a system of implicit equations for
and
, which
appear difficult to solve. In reality, the following simple recursion quickly converges to a solution,
see Figure 24.2.
Let us denote with and the location and dispersion parameters respectively obtained in the previous iteration. Convergence in the above routine occurs when the maximum relative norm [W] between two subsequent updates and are smaller than a required threshold
| | (24.25) |
Example 24.3. Consider as i.i.d. variables
two different zero swap rates (65.45), and exponential decay flexible probabilities
(23.6) with
half-life .
By setting ,
we apply maximum likelihood routine in Table 24.1, obtaining the MLFP estimates of location
(24.21) and dispersion (24.22)
| | (24.26) |
Example 24.4. Consider as i.i.d. variable
the
equity compounded return of Cisco Systems (CSCO), and exponential decay flexible probabilities
(23.6) with
half-life .
Since we are concerned of extreme outcomes we apply maximum likelihood
routine in Table 24.1 assuming a Cauchy distribution (49.164), namely
,
obtaining the MLFP estimates of location (24.21) and dispersion (24.22)
| | (24.27) |


