Untangling Complex Systems, page 99
tribution of exponential terms.
The term Pr( h) is the “prior probability,” and it codifies our knowledge about a possible spectrum
before collecting experimental data. It is expressed as
Pr( )
h
eS
∝ [B.6]
with S being the Information Entropy:
S( h) = − h( )log h(α) d
∫ α α [B.7]
Finally, the term Pr( h D) is the “posterior probability.” A single answer to the inverse Laplace
transform problem can be obtained by maximizing Pr( h D). The “posterior probability” is
S−1 2
χ
e 2
eQ
Pr( h D) =
=
[B.8]
ZlZD
ZlZD
Maximizing Pr( h D) means finding the maximum of the exponent Q = S − (1 2) 2
/ χ . Q is maximized
when the Entropy S is maximized whereas χ 2 is minimized. MEM becomes a “tug of war” between
maximizing the Entropy S and minimizing the χ 2 (see Figure B.3).
Mathematically, the Information Entropy has the important property that when maximizing it,
no possibilities are excluded. In other words, maximizing S in the absence of the experimental
data, we impose no correlations on the spectrum α( τ); we choose the spectrum that is minimally
structured or “maximally noncommittal.” This concept can be illustrated by a funny numerical
example: the so-called “Kangaroo problem” (see Figure B.4). The information available is that (1/3)
510
Appendix B
n 0
2
1
data datum
S( h) = ∫ [− h( τ)log h( τ)] dτ
χ 2 =
i − fiti
n 0
σ
data i=1
i
FIGURE B.3 The Maximum Entropy Method is a “tug of war” between the Information Entropy and the χ 2.
Left-handed
Left-handed
True
False
True
False
p
1
x
p 1
x
True
True
−
3
x
Blue-eyed
p
1
1
2
p 3
Blue-eyed
−
+ x
False
False
3
x
3
FIGURE B.4 The example of blue-eyed and left-handed kangaroos.
of kangaroos have blue eyes, and (1/3) are left-handed. The question is how many kangaroos are
both left-handed and blue-eyed. Of course, the two properties are unconnected (as far as we know,
unless one day we will discover that the genes responsible for blue-eyes are somehow correlated
with the genes responsible for left-handedness). We have four possibilities shown in the left Table of
Figure B.4. The normalization imposes that
px + p 1 + p 2 + p 3 = 1 [B.9]
We know that
1
px + p 1 =
3 [B.10]
1
px + p 2 =
3
If we fix px = x, p 1, and p 2 will be ( /
1 3 − x), and p 3 will be ( /
1 3 + x), as shown in the table on the
right of Figure B.4. According to the theory, the probability of two independent, uncorrelated events is given by the product of their single probabilities. In our case, it is expected that the probability
to have left-handed and blue-eyed kangaroos is ( /
1 )(
3 /
1 )
3 , i.e., ( /
1 )
9 . Is there some function of the
probabilities, pi, which, when maximized, yields the maximally noncommittal result? The answer
is yes, and it is just the Information Entropy function:
1
1
1
1
S = −
pilnpi = − xlnx +
− x ln
x
x ln
x
−
+
+
+
∑
2
[B.11]
3
3
3
3
i
If we determine the maximum of S with respect to x, we achieve x = 1/9.
Try by yourself, just as an exercise.
Maximizing other functions of pi imposes correlation between the eyes’ color and handedness
of kangaroos. For example, before the work by Shannon, other information measures had been
proposed. One of them was −∑ p 2 i. Such function does not give the same result of the Shannon
i
Entropy. In fact, maximizing −∑ p 2 i, we obtain x =1 1
/ 2, i.e., a negative correlation between the two
i
uncorrelated properties: blue-eyes and left-handedness.
Appendix B
511
Before ending this Appendix, it is important to present another consideration about the “likeli-
hood” function of equation [B.3]. When the experimental technique, chosen to collect the time-
resolved signal, taints the data with “normal” noise, the Gaussian distribution is the right function,
the nonlinear least-squares a good fitting procedure, and the minimization of χ 2 the right way to fit the
data. However, there are techniques that are better described by the Poisson statistics. An example
is the Time-Correlated Single Photon Counting technique. In such technique, the time scale is parti-
tioned in n channels, and the luminescence intensity decay curve is built by collecting the number of
independent photons emitted in each channel, c , , …, . Such counts follow the Poisson statistics.
1 c 2
cn
In statistics, there exist three probability distributions. The parent distribution is the Binomial one
(see Figure B.5) representing the number of successes in a sequence of n-independent yes/no experiments, each of which yields the successful outcome with probability p, and the negative result with
probability q = 1 − p. It is an asymmetric distribution unless p = 1/2. If we repeat the experiment n-times, the average value of the positive outcome is ν = np, and its standard deviation is σ = npq .
When the number of experiments is enormous, i.e., n → ∞, the Binomial distribution becomes symmet-
ric. It can be substituted by the Gaussian distribution (see Figure B.5) that is centered at ν = N, which is also the mean value; σ is its standard deviation that defines the width of the distribution. When the number of experiments is huge, i.e., n → ∞, and the probability of the positive result is tiny, i.e., p → 0, the Binomial distribution can be replaced by the Poisson distribution (presented in Figure B.5), wherein μ is its average value of ν , and σ = µ is the standard deviation. The Poisson distribution describes appropriately all those experiments where random events occurring with small probabilities p are considered.
Examples of such events are radioactive decays and the counting of emitted photons in a broad array of
channels. Usually, the Poisson distribution is asymmetric, but when µ → ∞, it becomes symmetric, and
it can be replaced by a Gaussian distribution having the same average and standard deviation.
In the case of the Single-Photon Counting Technique (Bajzer and Prendergast 1992), the prob-
ability of the observed counts c , p( c α τ ), follows the Poisson distribution:
i
i ; ( )
e−〈 ci〉 c ci
p( c
i
)
i ;α =
〈 〉 [B.12]
ci !
where 〈 ci〉 is the expected value of the number of counts in the i- th channel. Since the counts are
independent, the likelihood will be the joint probability
Pr( D h) = ∏ p( ci;α) [B.13]
i
The value 〈 ci〉 is modeled by the function R h for i = 1,…, n. The best estimates of the R h will i
i
be those values that maximize the likelihood and the entropy. Note that the log-likelihood func-
tion, ln Pr( D h)
, attains its maximum for the same value of h, as the likelihood function. It is
Binomial distribution
P( v, n) =
n!
pv (1 − p) n− v
v!( n − v)!
Gaussian
Poisson
distribution
distribution
n→∞
n→∞
p→0
e−( v − N)/2 σ 2
μ→∞
μv
p( v) =
1
p( v) = e− μ v!
σ√2 π
FIGURE B.5 Formula of the three distributions used in statistics to describe random events.
512
Appendix B
customary to determine the best h by minimizing −ln Pr( D h), which is equivalent to minimizing the Poisson deviance
D(α) = 2
c
∑{
}
iln c
i / Rhi c
− i − Rhi [B.14]
i
The Poisson deviance is close to the Gaussian χ 2 when the counts are very high. This result is in
agreement with the consideration that a Poisson distribution can be replaced by a Gaussian one
when the average value in each channel goes to infinity.
B.1 HINTS FOR FURTHER READING
Further information regarding the Maximum Entropy Method applied to time-resolved spectros-
copy, and image processing can be found on the website of Center for Molecular Modeling of the
National Institutes of Health (https://cmm.cit.nih.gov/maxent/). It is possible to download the software MemExp by P. J. Steinbach, for the analysis of the kinetic data.
Appendix C: Fourier Transform
of Waveforms
Computers are useless. They can only give you answers.
Pablo Picasso (1881–1973 AD)
Almost everything in the universe can be described as a waveform. Suffice to think that quantum
particles are wave packets. A waveform is a curve showing the shape of a wave as a function of time
or space or some other variables.
The Fourier Transform (Bracewell 1986) is a powerful mathematical tool that allows describing
any waveform as a sum of sinusoids, i.e., as a sum of sine and cosine functions.
Let us remember the general expression and features of a sinusoid, like a sine function:
y = A sin(2πω t +ϕ) [C.1]
The function y has a frequency ω that is the inverse of the time it takes to complete one period of
oscillation. The term A in equation [C.1] is the amplitude of y (see Figure C.1), and φ is its phase. In the bottom graph of Figure C.1, the phase difference (Δ φ) between two sinusoids, having the same frequency and the same amplitude, is highlighted.
When we encounter a waveform like that depicted in Figure C.2, we ask ourselves: “If we describe
the signal S as a sum of sinusoids having different frequencies, what frequencies are needed, and
what is the distribution of the amplitudes and phases as the functions of the frequency?” This ques-
tion has an answer because Fourier’s Theorem states that any signal of zero mean value can be
represented as a unique sum of sinusoids.
When we sample the signal S with a time resolution of ∆ t (see Figure C.2), and we collect N data, the total time window is TW = N * t
∆ . Inevitably, the smallest period we can infer from our set of
data is TMIN = 2 t
∆ . In other words, the largest frequency we can measure in our series of sinusoids
describing the data is ν MAX = 1 2
/ t
∆ . If the original signal S contains periods shorter than TMIN = 2 t
∆
or frequencies higher than ν MAX, our measurement is said to be aliased. If we want to evaluate short-
est periods, we must reduce the time interval ∆ t . The longest period we can obtain unambiguously
from a windowed trend is one that lasts the total sampling time, that is TMAX = N * t
∆ ; in other
words, the smallest frequency we can infer is ν MIN = 1/ N t
∆ . In addition, T defines our resolu-
MIN
tion in the determination of the period. Analogously, ν MIN defines the resolution in the frequency
determination. It derives that the maximum number of the sinusoids in the Fourier Transform will
be N/2. This statement makes sense because we have N data, and for the definition of each sinusoid
we need to determine two parameters: its amplitude and its phase. For each frequency ν = n/ N t
∆ ,
with n = 1, 2, …, N/2, its amplitude is
A
2
2
n =
an + bn [C.2]
and its phase is
b
ϕ
n
n = −arctan
[C.3]
n
a
where the coefficients, a and b , are given in terms of the original signal S( t) as
n
n
513
514
Appendix C
1.0
1/ ω
0.5
A
y 0.0
−0.5
−1.0
0.0
0.5
1.0
1.5
2.0
1.0
0.5
y 0.0
−0.5
Δ ϕ
−1.0
0.0
0.5
1.0
1.5
2.0
t
FIGURE C.1 Shape and features of the sinusoids.
4
Δ t
2
0
S
−2
−4
−6
Time
FIGURE C.2 An example of a waveform regarding the signal S as a function of time. Δ t represents the sampling period.
N
2
2π ni
N
2
2π ni
an =
S( i t
∑ ∆ )cos bn =
S( i t
∑ ∆ )sin
N
[C.4]
N
N
N
i=1
i=1
Expressions [C.4], [C.2], and [C.3] are referred to as the Discrete Fourier Transform (DFT).
The original set of data S( iΔ t) is finally expressed in terms of the sum of the N/2 sinusoids: A
N /2
S( i t
∆ ) = 0 +
An cos(2πν ni t
∆ +ϕ n) [C.5]
2
∑ n=1
Appendix C
515
Frequency
400 0.0
0.1
0.2
0.3
0.4
0.5
200
Phase
0
1.2
1.0
0.8
0.6
Amplitude 0.4
0.2
0.0 0.0
0.1
0.2
0.3
0.4
0.5
Frequency
FIGURE C.3 Discrete Fourier Transform of the data plotted in Figure C.2.
Equation [C.5] is called the Inverse Discrete Fourier Transform. The term A 0 2
/ represents any mean
value of the signal which is not accounted for in the other sinusoids. It comes from the definitions of
a and b (cf. equation [C.4]) when n
n
n
= 0, which corresponds to frequency 0.
When we decompose our signal S into its component sinusoids, the resulting graph reporting
amplitude or phase as a function of frequency is referred to as its Fourier spectrum. In Figure C.3, we report the Fourier spectrum of the 53 data plotted in Figure C.2. It consists of 27 frequencies included between 0 and 0.5.
Usually, the DFT is computed by an algorithm known as the Fast Fourier Transform (FFT). The
key advantage of the FFT over the DFT is that the operational complexity decreases from O(N2) for
a DFT to O(Nlog (N)) for the FFT.
2
Appendix D: Errors and Uncertainties
in Laboratory Experiments
Those sciences are vain and full of errors which are not born from experiment, the mother of
certainty.
Leonardo da Vinci (1452–1519 AD)
Science is said to be exact not because it is based on infinitely exact assertions but because its rigor-
ous methodology allows estimating the extent of the uncertainty associated with any determination.
D.1 ERRORS IN DIRECT MEASUREMENTS
Many scientific assertions are grounded in experiments. Leonardo da Vinci was used to say that
scientific knowledge without experiments is vain and full of mistakes. However, we are aware that
even when we perform experiments, we make errors.
Usually, an experiment involves four protagonists: the observer raising a question and looking
for an answer; the object or the phenomenon that is inquired; an instrument, and the surrounding
environment (see Figure D.1).
The Galilean experimental method requires that the object or the phenomenon under study is ana-
