Question

1. (3 points) In the lectures, we discussed that px,y characterizes how much two random variables X

and Y are linearly dependent on each other. In particular, the larger px,y, the more useful X is in

estimating the value of Y as an affine function of X, and vice versa. In this question, we will investigate

the implication of px,y taking its maximum value, i.e., px,y = 1, or -1.

For simplicity, suppose that X and Y have zero means. Consider estimating the value of X as a linear

function of Y, i.e., aY with some constant a.

(a) Consider E[(X-aY)2] for an arbitrary real constant a. This quantity is the expected estimation

error if we consider estimating the value of X as aY. It is easy to see that E[(X-aY)²] should be

non-negative for any value of a. Find the value of a, denoted by a*, that minimizes E[(X-GY)²].

(Hint: E[(X-aY)²] can be written as a second order polynomial of a with its coefficients expressed

in terms of the variance and covariance of X and Y.)

(b) With a* found earlier, E[(X - a*Y)²] should be still non-negative. Using this fact to prove that

PX, Y| ≤ 1.

(c) Based on the derivations you made for earlier questions, explain what "px,y|² 1" implies

regarding the relation between X and Y. (Hint: For a random variable Z, if E[Z²] = 0, then

Z=0 with probability one.)

Question image 1