Download as pdf or txt
Download as pdf or txt
You are on page 1of 62

Generalized Additive Models: An

Introduction with R (Chapman &


Hall/CRC Texts in Statistical Science)
1st Edition ■ Ebook PDF Version
Visit to download the full and correct content document:
https://ebookmass.com/product/generalized-additive-models-an-introduction-with-r-ch
apman-hall-crc-texts-in-statistical-science-1st-edition-ebook-pdf-version/
Generalized Additive Models:
an introduction with R
COPYRIGHT CRC DO NOT DISTRIBUTE

Simon N. Wood
viii CONTENTS
4.9.1 Single smooths 192
4.9.2 GAMs and their components 195
4.9.3 Unconditional Bayesian confidence intervals 198
4.10 Further GAM theory 200
4.10.1 Comparing GAMs by hypothesis testing 200
4.10.2 ANOVA decompositions and Nesting 202
4.10.3 The geometry of penalized regression 204
4.10.4 The “natural” parameterization of a penalized smoother 205
4.11 Other approaches to GAMs 208
4.11.1 Backfitting GAMs 209
4.11.2 Generalized smoothing splines 211
4.12 Exercises 213

5 GAMs in practice: mgcv 217


5.1 Cherry trees again 217
5.1.1 Finer control of gam 219
5.1.2 Smooths of several variables 221
5.1.3 Parametric model terms 224
5.2 Brain Imaging Example 226
5.2.1 Preliminary Modelling 228
5.2.2 Would an additive structure be better? 232
5.2.3 Isotropic or tensor product smooths? 233
5.2.4 Detecting symmetry (with by variables) 235
5.2.5 Comparing two surfaces 237
5.2.6 Prediction with predict.gam 239
Prediction with lpmatrix 241
5.2.7 Variances of non-linear functions of the fitted model 242
5.3 Air Pollution in Chicago Example 243
5.4 Mackerel egg survey example 249
5.4.1 Model development 250
5.4.2 Model predictions 255
CONTENTS ix
5.5 Portuguese larks 257
5.6 Other packages 261
5.6.1 Package gam 261
5.6.2 Package gss 263
5.7 Exercises 265

6 Mixed models: GAMMs 273


6.1 Mixed models for balanced data 273
6.1.1 A motivating example 273
The wrong approach: a fixed effects linear model 274
The right approach: a mixed effects model 276
6.1.2 General principles 277
6.1.3 A single random factor 278
6.1.4 A model with two factors 281
6.1.5 Discussion 286
6.2 Linear mixed models in general 287
6.2.1 Estimation of linear mixed models 288
6.2.2 Directly maximizing a mixed model likelihood in R 289
6.2.3 Inference with linear mixed models 290
Fixed effects 290
Inference about the random effects 291
6.2.4 Predicting the random effects 292
6.2.5 REML 293
The explicit form of the REML criterion 295
6.2.6 A link with penalized regression 296
6.3 Linear mixed models in R 297
6.3.1 Tree Growth: an example using lme 298
6.3.2 Several levels of nesting 303
6.4 Generalized linear mixed models 303
6.5 GLMMs with R 305
6.6 Generalized Additive Mixed Models 309
x CONTENTS
6.6.1 Smooths as mixed model components 309
6.6.2 Inference with GAMMs 311
6.7 GAMMs with R 312
6.7.1 A GAMM for sole eggs 312
6.7.2 The Temperature in Cairo 314
6.8 Exercises 318

A Some Matrix Algebra 325


A.1 Basic computational efficiency 325
A.2 Covariance matrices 326
A.3 Differentiating a matrix inverse 326
A.4 Kronecker product 327
A.5 Orthogonal matrices and Householder matrices 327
A.6 QR decomposition 328
A.7 Choleski decomposition 328
A.8 Eigen-decomposition 329
A.9 Singular value decomposition 330
A.10 Pivoting 331
A.11 Lanczos iteration 331

B Solutions to exercises 335


B.1 Chapter 1 335
B.2 Chapter 2 340
B.3 Chapter 3 345
B.4 Chapter 4 347
B.5 Chapter 5 354
B.6 Chapter 6 363

Bibliography 373

Index 378
Preface

This book is designed for readers wanting a compact, but thorough, introduction to
linear models, generalized linear models , generalized additive models, and the mixed
model extension of these, with particular emphasis on generalized additive models.
The aim is to provide a full, but concise, theoretical treatment, explaining how the
models and methods work, in order to underpin quite extensive material on practical
application of the models using R.

Linear models are statistical models in which a univariate response is modelled as the
sum of a ‘linear predictor’ and a zero mean random error term. The linear predictor
depends on some predictor variables, measured with the response variable, and some
unknown parameters, which must be estimated. A key feature of linear models is
that the linear predictor depends linearly on these parameters. Statistical inference
with such models is usually based on the assumption that the response variable has
a normal distribution. Linear models are used widely in most branches of science,
both in the analysis of designed experiments, and for other modeling tasks, such as
polynomial regression. The linearity of the models endows them with some rather
elegant theory, which is explored in some depth in Chapter 1, alongside practical
examples of their use.

Generalized linear models (GLMs) somewhat relax the strict linearity assumption of
linear models, by allowing the expected value of the response to depend on a smooth
monotonic function of the linear predictor. Similarly the assumption that the response
is normally distributed is relaxed by allowing it to follow any distribution from the
exponential family (for example, normal, Poisson, binomial, gamma etc.). Inference
for GLMs is based on likelihood theory, as is explained, quite fully, in chapter 2,
where the practical use of these models is also covered.

A Generalized Additive Model (GAM) is a GLM in which part of the linear pre-
dictor is specified in terms of a sum of smooth functions of predictor variables. The
exact parametric form of these functions is unknown, as is the degree of smoothness
appropriate for each of them. To use GAMs in practice requires some extensions to
GLM methods:

1. The smooth functions must be represented somehow.


2. The degree of smoothness of the functions must be made controllable, so that
models with varying degrees of smoothness can be explored.

xi
xii PREFACE
3. Some means for estimating the most appropriate degree of smoothness from data
is required, if the models are to be useful for more than purely exploratory work.
This book provides an introduction to the framework for Generalized Additive Mod-
elling in which (i) is addressed using basis expansions of the smooth functions, (ii) is
addressed by estimating models by penalized likelihood maximization, in which wig-
gly models are penalized more heavily than smooth models in a controllable manner,
and (iii) is performed using methods based on cross validation or sometimes AIC or
Mallow’s statistic. Chapter 3 introduces this framework, chapter 4 provides details
of the theory and methods for using it, and chapter 5 illustrated the practical use of
GAMs using the R package mgcv.
The final chapter of the book looks at mixed model extensions of linear, general-
ized linear, and generalized additive models. In mixed models, some of the unknown
coefficients (or functions) in the model linear predictor are now treated as random
variables (or functions). These ‘random effects’ are treated as having a covariance
structure that itself depends on some unknown fixed parameters. This approach en-
ables the use of more complex models for the random component of data, thereby
improving our ability to model correlated data. Again theory and practical applica-
tion are presented side by side.
I assume that most people are interested in statistical models in order to use them,
rather than to gaze upon the mathematical beauty of their structure, and for this rea-
son I have tried to keep this book practically focused. However, I think that practical
work tends to progress much more smoothly if it is based on solid understanding of
how the models and methods used actually work. For this reason, the book includes
fairly full explanations of the theory underlying the methods, including the underly-
ing geometry, where this is helpful. Given that the applied modelling involves using
computer programs, the book includes a good deal of material on statistical mod-
elling in R. This approach is now fairly standard when writing about practical sta-
tistical analysis, but in addition Chapter 3 attempts to introduce GAMs by having
the reader ‘build their own’ GAM using R: I hope that this provides a useful way of
quickly gaining a rather solid familiarity with the fundamentals of the GAM frame-
work presented in this book. Once the basic framework is mastered from chapter 3,
the theory in chapter 4 is really filling in details, albeit practically important ones.
The book includes a moderately high proportion of practical examples which re-
flect the reality that statistical modelling problems are usually quite involved, and
rarely require only straightforward brain-free application of some standard model.
This means that some of the examples are fairly lengthy, but do provide illustration
of the process of producing practical models of some scientific utility, and of check-
ing those models. They also provide much greater scope for the reader to decide that
what I’ve done is utter rubbish.
Working through this book from Linear Models, through GLMs to GAMs and even-
tually GAMMs, it is striking that as model flexibility increases, so that the models
become better able to describe the reality that we believe generated a set of data, so
the methods for inference become less well founded. The linear model class is quite
PREFACE xiii
restricted, but within it, hypothesis testing and interval estimation are exact, while
estimation is unbiased. For the larger class of GLMs this exactness is generally lost
in favour of the large sample approximations of general likelihood theory, while esti-
mators themselves are consistent, but not necessarily unbiased. Generalizing further
to GAMs, penalization lowers the convergence rates of estimators, hypothesis testing
is only approximate, and satisfactory interval estimation seems to require the adop-
tion of a Bayesian approach. With time, improved theory will hopefully reduce these
differences. In the meantime, this book is offered in the belief that it is usually better
to be able to say something approximate about the right model, rather than something
very precise about the wrong model.
Life is too short to spend too much of it reading statistics text books. This book is of
course an exception to this rule and should be read from cover to cover. However, if
you don’t feel inclined to follow this suggestion, here are some alternatives.
• For those who are already very familiar with linear models and GLMs, but want
to use GAMs with a reasonable degree of understanding: work through Chapter 3
and read chapter 5, trying some exercises from both, use chapter 4 for reference.
Perhaps skim the other chapters.
• For those who want to focus only on practical modelling in R, rather than theory.
Work through the following: 1.5, 1.6.4, 1.7, 2.3, Chapter 5, 6.3, 6.5 and 6.7.
• For those familiar with the basic idea of setting up a GAM using basis expansions
and penalties, but wanting to know more about the underlying theory and practical
application: work through Chapters 4 and 5, and probably 6.
• For those not interested in GAMs, but wanting to know about linear models,
GLMs and mixed models. Work through Chapters 1 and 2, and Chapter 6 up
to section 6.6.
The book is written to be accessible to numerate researchers and students from the
last two years of an undergraduate programme upwards. It is designed to be used ei-
ther for self study, or as the text for the ‘regression modelling’ strand of mathematics
and/or statistics degree programme. Some familiarity with statistical concepts is as-
sumed, particularly with notions such as random variable, expectation, variance and
distribution. Some knowledge of matrix algebra is also needed, although Appendix
A is intended to provide all that is needed beyond the fundamentals.
Finally, I’d like to thank the people who have in various ways helped me out in the
writing of this book, or in the work that lead to writing it. Among these, are Lucy
Augustin, Nicole Augustin, Miguel Bernal, Steve Blythe, David Borchers, Mark
Bravington, Steve Buckland, Richard Cormack, José Pedro Granadeiro ,Chong Gu,
Bill Gurney, John Halley, Joe Horwood, Sharon Hedley, Peter Jupp, Alain Le Tetre,
Stephan Lang, Mike Lonergan, Henric Nilsson, Roger D. Peng, Charles Paxton,
Björn Stollenwerk, Yorgos Stratoudaki, the R core team in particular Kurt Hornik
and Brian Ripley, the Little Italians and the RiederAlpinists. I am also very grateful
to the people who have sent me bug reports and suggestions which have greatly im-
proved the the mgcv package over the last few years: the list is rather too long to
reproduce here, but thankyou.
CHAPTER 1

Linear Models

How old is the universe? The standard big-bang model of the origin of the universe
says that it expands uniformly, and locally, according to Hubble’s law:
y = βx
where y is the relative velocity of any two galaxies separated by distance x, and β is
“Hubble’s constant”(in standard astrophysical notation y ≡ v, x ≡ d and β ≡ H0 ).
β −1 gives the approximate age of the universe, but β is unknown and must somehow
be estimated from observations of y and x, made for a variety of galaxies at different
distances from us.
Figure 1.1 plots velocity against distance for 24 galaxies, according to measurements
made using the Hubble Space Telescope. Velocities are assessed by measuring the
Doppler effect red shift in the spectrum of light observed from the Galaxies con-
cerned, although some correction for ‘local’ velocity components is required. Dis-
1500
Velocity (kms−1)
1000
500

5 10 15 20
Distance (Mpc)

Figure 1.1 A Hubble diagram showing the relationship between distance, x, and velocity, y,
for 24 Galaxies containing Cepheid stars. The data are from the Hubble Space Telescope key
project to measure the Hubble constant as reported in Freedman et al. (2001).

1
2 LINEAR MODELS
tance measurement is much less direct, and is based on the 1912 discovery, by Hen-
rietta Leavit, of a relationship between the period of a certain class of variable stars,
known as the Cepheids, and their luminosity. The intensity of Cepheids varies regu-
larly with a period of between 1.5 and something over 50 days, and the mean intensity
increases predictably with period. This means that, if you can find a Cepheid, you can
tell how far away it is, by comparing its apparent brightness to its period predicted
intensity.
It is clear, from the figure, that the observed data do not follow Hubble’s law exactly,
but given the measurement process, it would be surprising if they did. Given the
apparent variability, what can be inferred from these data? In particular: (i) what
value of β is most consistent with the data? (ii) what range of β values is consistent
with the data and (iii) are some particular theoretically derived values of β consistent
with the data? Statistics is about trying to answer these three sorts of question.
One way to proceed is to formulate a linear statistical model of the way that the data
were generated, and to use this as the basis for inference. Specifically, suppose that,
rather than being governed directly by Hubble’s law, the observed velocity is given
by Hubble’s constant multiplied by the observed distance plus a ‘random variability’
term. That is
yi = βxi + i i = 1 . . . 24 (1.1)
where the i terms are independent random variables such that E(i ) = 0 and
E(2i ) = σ 2 . The random component of the model is intended to capture the fact
that if we gathered a replicate set of data, for a new set of galaxies, Hubble’s law
would not change, but the apparent random variation from it would be different, as a
result of different measurement errors. Notice that it is not implied that these errors
are completely unpredictable: their mean and variance are assumed to be fixed, it is
only their particular values, for any particular galaxy, that are not known.

1.1 A simple linear model

This section develops statistical methods for a simple linear model of the form (1.1).
This allows the key concepts of linear modelling to be introduced without the dis-
traction of any mathematical difficulty.
Formally, consider n observations, xi , yi , where yi is an observation on random vari-
able, Yi , with expectation, µi ≡ E(Yi ). Suppose that an appropriate model for the
relationship between x and y is:
Yi = µi + i where µi = xi β. (1.2)
Here β is an unknown parameter and the i are mutually independent zero mean
random variables, each with the same variance σ 2 . So the model says that Y is given
by x multiplied by a constant plus a random term. Y is an example of a response
variable, while x is an example of a predictor variable. Figure 1.2 illustrates this
model for a case where n = 8.
A SIMPLE LINEAR MODEL 3

Yi
εi

βx

xi

Figure 1.2 Schematic illustration of a simple linear model with one explanatory variable.

Simple least squares estimation

How can β, in model (1.2) be estimated from the xi , yi data? A sensible approach
is to choose a value of β that makes the model fit closely to the data. To do this we
need to define a measure of how well, or how badly, a model with a particular β fits
the data. One possible measure is the residual sum of squares of the model:
n
X n
X
S= (yi − µi )2 = (yi − xi β)2
i=1 i=1

If we have chosen a good value of β, close to the true value, then the model pre-
dicted µi should be relatively close to the yi , so that S should be small, whereas poor
choices will lead to µi far from their corresponding yi , and high values of S. Hence
β can be estimated by minimizing S w.r.t. β and this is known as the method of least
squares.
To minimize S, differentiate w.r.t. β:
n
∂S X
=− 2xi (yi − xi β)
∂β i=1

and set the result to zero to find β̂, the least squares estimate of β:
n
X n
X n
X n
X n
X
− 2xi (yi − xi β̂) = 0 ⇒ xi yi − β̂ x2i = 0 ⇒ β̂ = xi yi / x2i .∗
i=1 i=1 i=1 i=1 i=1

1.1.1 Sampling properties of β̂

To evaluate the reliability of the least squares estimate, β̂, it is useful to consider the
sampling properties of β̂. That is, we should consider some properties of the distri-
bution of β̂ values, which would be obtained from repeated independent replication
∗ ∂ 2 S/∂β 2 = 2 P x2 which is clearly positive, so a minimum of S has been found.
i
4 LINEAR MODELS
of the xi , yi data used for estimation. To do this, it is helpful to introduce the concept
of an estimator, which is obtained by replacing the observations, yi , in the estimate
of β̂ by the random variables, Yi , to obtain
n
X n
X
β̂ = xi Yi / x2i .
i=1 i=1

Clearly the estimator, β̂, is a random variable and we can therefore discuss its distri-
bution. For now, consider only the first two moments of that distribution.

The expected value of β̂ is obtained as follows:


n n n n n n
!
X X X X X X
E(β̂) = E xi Yi / x2i = xi E(Yi )/ x2i = x2i β/ x2i = β.
i=1 i=1 i=1 i=1 i=1 i=1

So β̂ is an unbiased estimator — its expected value is equal to the true value of the
parameter that it is supposed to estimate.

Unbiasedness is a reassuring property, but knowing that an estimator gets it right on


average, does not tell us much about how good any one particular estimate is likely to
be: for this we also need to know how much estimates would vary from one replicate
data set to the next — we need to know the estimator variance.

From general probability theory we know that if Y1 , Y2 , . . . , Yn are independent ran-


dom variables and a1 , a2 , . . . an are real constants then
!
X X
var ai Yi = a2i var(Yi ).
i i

But we can write


X X
β̂ = ai Yi where ai = xi / x2i ,
i i

and from the original model specification we have that var(Yi ) = σ 2 for all i. Hence,
!2 !−1
X X X
2 2 2 2
var(β̂) = xi / xi σ = xi σ2 . (1.3)
i i i

In most circumstances σ 2 itself is an unknown parameter and must also be estimated.


Since σ 2 is the variance of the i , it makes sense to estimate it using the variance of
the estimated i , the model residuals, ˆi = yi − xi β̂. An unbiased estimator of σ 2 is:
1 X
σˆ2 = (yi − xi β̂)2
n−1 i

(proof of unbiasedness is given later for the general case). Plugging this into (1.3)
obviously gives an unbiased estimate of the variance of β̂.
A SIMPLE LINEAR MODEL 5
1.1.2 So how old is the universe?

The least squares calculations derived above are available as part of the statistical
package and environment R. The function lm fits linear models to data, including
the simple example currently under consideration. The Cepheid distance — velocity
data shown in figure 1.1 are stored in a data frame† hubble. The following R code
fits the model and produces the (edited) output shown.

> data(hubble)
> hub.mod <- lm(y˜x-1,data=hubble)
> summary(hub.mod)

Call:
lm(formula = y ˜ x - 1, data = hubble)

Coefficients:
Estimate Std. Error
x 76.581 3.965

The call to lm passed two arguments to the function. The first is a model formula,
y˜x-1, specifying the model to be fitted: the name of the response variable is to
the left of ‘˜’ while the predictor variable is specified on the right; the ‘-1’ term
indicates that the model has no ‘intercept’ term, i.e. that the model is a straight line
through the origin. The second (optional) argument gives the name of the data frame
in which the variables are to be found. lm takes this information and uses it to fit the
model by least squares: the results are returned in a ‘fitted model object’, which in
this case has been assigned to an object called hub.mod for later examination. ‘<-’
is the assignment operator, and hub.mod is created by this assignment (overwriting
any previously existing object of this name).
The summary function is then used to examine the fitted model object. Only part
of its output is shown here: β̂ and the estimate of the standard error of β̂ (the square
root of the estimated variance of β̂, derived above). Before using these quantities
it is important to check the model assumptions. In particular we should check the
plausibility of the assumptions that the i are independent and all have the same
variance. The way to do this is to examine residual plots.
The ‘fitted values’ of the model are defined as µ̂i = β̂xi , while the residuals are
simply ˆi = yi − µ̂i . A plot of residuals against fitted values is often useful and the
following produces the plot in figure 1.3 (a).

plot(fitted(hub.mod),residuals(hub.mod),xlab="fitted values",
ylab="residuals")

What we would like to see, in such a plot, is an apparently random scatter of residu-
als around zero, with no trend in either the mean of the residuals, or their variability,

† A data frame is just a two dimensional array of data in which the values of different variables are stored
in different named columns.
6 LINEAR MODELS

600

200
400
a b

200

100
residuals

residuals
0

0
−200

−100
−600

−300
500 1000 1500 500 1000 1500

fitted values fitted values

Figure 1.3 Residuals against fitted values for (a) the model (1.1) fitted to all the data in figure
1.1 and (b) the same model fitted to data with two substantial outliers omitted.

as the fitted values increase. A trend in the mean violates the independence assump-
tion, and is usually indicative of something missing in the model structure, while a
trend in the variability violates the constant variance assumption. The main problem-
atic feature of figure 1.3(a) is the presence of two points with very large magnitude
residuals, suggesting a problem with the constant variance assumption. It is proba-
bly prudent to repeat the model fit, with and without these points, to check that they
are not having undue influence on our conclusions. The following code omits the
offending points and produces a new residual plot shown in figure 1.3(b).

> hub.mod1 <- lm(y˜x-1,data=hubble[-c(3,15),])


> summary(hub.mod1)

Call:
lm(formula = y ˜ x - 1, data = hubble[-c(3, 15), ])

Coefficients:
Estimate Std. Error
x 77.67 2.97

> plot(fitted(hub.mod1),residuals(hub.mod1),
+ xlab="fitted values",ylab="residuals")

The omission of the two large outliers has improved the residuals and changed β̂
somewhat, but not drastically.

The Hubble constant estimates have units of (km)s−1 (Mpc)−1 . A Mega-parsec is


3.09 × 1019 km, so we need to divide β̂ by this amount, in order to obtain Hubble’s
constant with units of s−1 . The approximate age of the universe, in seconds, is then
given by the reciprocal of β̂. Here are the two possible estimates expressed in years:
A SIMPLE LINEAR MODEL 7
> hubble.const <- c(coef(hub.mod),coef(hub.mod1))/3.09e19
> age <- 1/hubble.const
> age/(60ˆ2*24*365)
12794692825 12614854757

Both fits give an age of around 13 billion years. So we now have an idea of the best
estimate of the age of the universe, but what range of ages would be consistent with
the data?

1.1.3 Adding a distributional assumption

So far everything done with the simple model has been based only on the model
equations and the two assumptions of independence and equal variance, for the re-
sponse variable. If we wish to go further, and find confidence intervals for β, or test
hypotheses related to the model, then a further distributional assumption will be nec-
essary.
Specifically, assume that i ∼ N (0, σ 2 ) for all i, which is equivalent to assuming
Yi ∼ N (xi β, σ 2 ). Now we have already seen that β̂ is just a weighted sum of Yi ,
but the Yi are now assumed to be normal random variables, and a weighted sum of
normal random variables is itself a normal random variable. Hence the estimator, β̂,
must be a normal random variable. Since we have already established the mean and
variance of β̂, we have that
 X  
−1
β̂ ∼ N β, xi σ2 . (1.4)

Testing hypotheses about β

One thing we might want to do is to try and evaluate the consistency of some hy-
pothesized value of β with the data. For example some Creation Scientists estimate
the age of the universe to be 6000 years, based on a reading of the Bible. This would
imply that β = 163 × 106‡ . The consistency with data of such a hypothesized value
for β, can be based on the probability that we would have observed the β̂ actually
obtained, if the true value of β was the hypothetical one.
Specifically, we can test the null hypothesis, H0 : β = β0 , versus the alternative
hypothesis, H1 : β 6= β0 , for some specified value β0 , by examining the probability
of getting the observed β̂, or one further from β0 , assuming H0 to be true. If σ 2 were
known then we could work directly from (1.4), as follows.
The probability required is known as the p-value of the test. It is the probability of
getting a value of β̂ at least as favourable to H1 as the one actually observed, if H0 is

‡ This isn’t really valid, of course, since the Creation Scientists are not postulating a big bang theory.
8 LINEAR MODELS
§
actually true . In this case it helps to distinguish notationally between the estimate,
β̂obs , and estimator β̂. The p-value is then
h i
p = Pr |β̂ − β0 | ≥ |β̂obs − β0 | H0
" #
|β̂ − β0 | |β̂obs − β0 |
= ≥ H0
σβ̂ σβ̂
= Pr[|Z| > |z|]
where Z ∼ N (0, 1), z = (β̂obs − β0 )/σβ̂ and σβ̂ = ( x2i )−1 σ 2 . Hence, having
P
formed z, the p-value is easily worked out, using the cumulative distribution function
for the standard normal, built into any statistics package. Small p-values suggest that
the data are inconsistent with H0 , while large values indicate consistency. 0.05 is
often used as the boundary between ‘small’ and ‘large’ in this context.
In reality σ 2 is usually unknown. Broadly the same testing procedure can still be
adopted, by replacing σ with σ̂, but we need to somehow allow for the extra uncer-
tainty that this introduces (unless the sample size is very large). It turns out that if
H0 : β = β0 is true then
β̂ − β0
T ≡ ∼ tn−1
σ̂β̂
where n is the sample size, σ̂β̂ = ( x2i )−1 σ̂, and tn−1 is the t-distribution with
P
n − 1 degrees of freedom. This result is proven in section 1.3. It is clear that large
magnitude values of T favour H1 , so using T as the test statistic, in place of β̂ we
can calculate a p-value by evaluating
p = Pr[|T | > |t|]
where T ∼ tn−1 and t = (β̂obs − β0 )/σ̂β̂ . This is easily evaluated using the c.d.f.
of the t distributions, built into any decent statistics package. Here is some code to
evaluate the p-value for H0 : the Hubble constant is 163000000.

> cs.hubble <- 163000000


> t.stat<-(coef(hub.mod1)-cs.hubble)/
+ summary(hub.mod1)$coefficients[2]
> pt(t.stat,df=21)*2 # 2 because of |T| in p-value defn.
3.906388e-150

i.e. as judged by the test statistic, t, the data would be hugely improbable if β =
1.63 × 108 . It would seem that the hypothesized value can be rejected rather firmly
(in this case, using the data with the outliers increases the p-value by a factor of 1000
or so).
Hypothesis testing is particularly useful when there are good reasons to want to stick
with some null hypothesis, until there is good reason to reject it. This is often the

§ This definition holds for any hypothesis test, if the specific ‘β̂’ is replaced by the general ‘a test statis-
tic’.
A SIMPLE LINEAR MODEL 9
case when comparing models of differing complexity: it is often a good idea to retain
the simpler model until there is quite firm evidence that it is inadequate. Note one
interesting property of hypothesis testing. If we choose to reject a null hypothesis
whenever the p-value is less than some fixed level, α (often termed the significance
level of a test), then we will inevitably reject a proportion, α, of correct null hypothe-
ses. We could try and reduce the probability of such mistakes by making α very
small, but in that case we pay the price of reducing the probability of rejecting H0
when it is false!

Confidence intervals

Having seen how to test whether a particular hypothesized value of β is consistent


with the data, the question naturally arises of what range of values of β would be
consistent with the data. To do this, we need to select a definition of ‘consistent’: a
common choice is to say that any parameter value is consistent with the data if it
results in a p-value of ≥ 0.05, when used as the null value in a hypothesis test.
Sticking with the Hubble constant example, and working at a significance level of
0.05, we would have rejected any hypothesized value for the constant, that resulted
in a t value outside the range (−2.08, 2.08), since these values would result in p-
values of less than 0.05. The R function qt can be used to find such ranges: e.g.
qt(c(0.025,0.975),df=21) returns the range of the middle 95% of t21 ran-
dom variables. So we would accept any β0 fulfilling:
β̂ − β0
−2.08 ≤ ≤ 2.08
σ̂β̂
which re-arranges to give the interval
β̂ − 2.08σ̂β̂ ≤ β0 ≤ β̂ + 2.08σ̂β̂ .
Such an interval is known as a ‘95% Confidence interval’ for β.
The defining property of a 95% confidence interval is this: if we were to gather an
infinite sequence of independent replicate data sets, and calculate 95% confidence
intervals for β from each, then 95% of these intervals would include the true β, and
5% would not. It is easy to see how this comes about. By construction, a hypothesis
test with a significance level of 5%, rejects the correct null hypothesis for 5% of
replicate data sets, and accepts it for the other 95% of replicates. Hence 5% of 95%
confidence intervals must exclude the true parameter, while 95% include it.
For the Hubble example, a 95% CI for the constant (in the usual astro-physicists
units) is given by:

> sigb <- summary(hub.mod1)$coefficients[2]


> h.ci<-coef(hub.mod1)+qt(c(0.025,0.975),df=21)*sigb
> h.ci
[1] 71.49588 83.84995
10 LINEAR MODELS
This can be converted to a confidence interval for the age of the universe, in years, as
follows:

> h.ci<-h.ci*60ˆ2*24*365.25/3.09e19 # convert to 1/years


> sort(1/h.ci)
[1] 11677548698 13695361072

i.e. the 95% CI is (11.7,13.7) billion years. Actually this ‘Hubble age’ is the age
of the universe if it has been expanding freely, essentially unfettered by gravitation.
If the universe is really ‘matter dominated’ then the galaxies should be slowed by
gravity over time so that the universe would actually be younger than this, but it is
time to get on with the subject of this book.

1.2 Linear models in general

The simple linear model, introduced above, can be generalized by allowing the re-
sponse variable to depend on multiple predictor variables (plus an additive constant).
These extra predictor variables can themselves be transformations of the original pre-
dictors. Here are some examples, for each of which a response variable datum, yi ,
is treated as an observation on a random variable, Yi , where E(Yi ) ≡ µi , the i are
zero mean random variables, and the βj are model parameters, the values of which
are unknown and will need to be estimated using data.

1. µi = β0 + xi β1 , Yi = µi + i , is a straight line relationship between y and


predictor variable, x.
2. µi = β0 + xi β1 + x2i β2 + x3i β3 , Yi = µi + i , is a cubic model of the relationship
between y and x.
3. µi = β0 + xi β1 + zi β2 + log(xi zi )β3 , Yi = µi + i , is a model in which y
depends on predictor variables x and z and on the log of their product.

Each of these is a linear model because the i terms and the model parameters, βj ,
enter the model in a linear way. Notice that the predictor variables can enter the model
non-linearly. Exactly as for the simple model, the parameters of these models can be
estimated by finding the βjP values which make the models best fit the observed data,
in the sense of minimizing i (yi − µi )2 . The theory for doing this will be developed
in section 1.3, and that development is based entirely on re-writing the linear model
using using matrices and vectors.

To see how this re-writing is done, consider the straight line model given above.
Writing out the µi equations for all n pairs, (xi , yi ), results in a large system of
LINEAR MODELS IN GENERAL 11
linear equations:
µ1 = β 0 + x 1 β1
µ2 = β 0 + x 2 β1
µ3 = β 0 + x 3 β1
. .
. .
µn = β0 + xn β1

which can be re-written in matrix-vector form as


   
µ1 1 x1
 µ2   1 x2 
    
 µ3   1 x3  β0
 =  .
 .   . . 
    β1
 .   . . 
µn 1 xn
So the model has the general form µ = Xβ, i.e. the expected value vector µ is given
by a model matrix (also known as a design matrix), X, multiplied by a parameter
vector, β. All linear models can be written in this general form.
As a second illustration, the cubic example, given above, can be written in matrix
vector form as
1 x1 x21 x31 
   
µ1
 µ2   1 x2 x22 x32  β0

 µ3   1 x3 x23 x33   β1 
   
 . = .
    .
   . . .  β2
 
 .   . . . .  β3
µn 1 xn x2n x3n
Models in which data are divided into different groups, each of which are assumed
to have a different mean, are less obviously of the form µ = Xβ, but in fact they
can be written in this way, by use of dummy indicator variables. Again, this is most
easily seen by example. Consider the model
µi = βj if observation i is in group j,
and suppose that there are three groups, each with 2 data. Then the model can be
re-written    
µ1 1 0 0
 µ2   1 0 0   
 µ3   0 1 0  β0
   
 µ4  =  0 1 0  β1 .
    
    β2
 µ5   0 0 1 
µ6 0 0 1
Variables indicating the group to which a response observation belongs, are known as
factor variables. Somewhat confusingly, the groups themselves are known as levels
12 LINEAR MODELS
of a factor. So the above model involves one factor, ‘group’, with three levels. Mod-
els of this type, involving factors, are commonly used for the analysis of designed
experiments. In this case the model matrix depends on the design of the experiment
(i.e on which units belong to which groups), and for this reason the terms ‘design
matrix’ and ‘model matrix’ are often used interchangeably. Whatever it is called ,X
is absolutely central to understanding the theory of linear models, generalized linear
models and generalized additive models.

1.3 The theory of linear models

This section shows how the parameters, β, of the linear model


µ = Xβ, y ∼ N (µ, In σ 2 )
can be estimated by least squares. It is assumed that X is a matrix, with n rows
and p columns. It will be shown that the resulting estimator, β̂, is unbiased, has the
lowest variance of any possible linear estimator of β, and that, given the normality
of the data, β̂ ∼ N (β, (XT X)−1 σ 2 ). Results are also derived for setting confidence
limits on parameters and for testing hypotheses about parameters — in particular the
hypothesis that several elements of β are simultaneously zero.
In this section it is important not to confuse the length
√ of a vector with its dimension.
For example (1, 1, 1)T has dimension 3 and length 3. Also note that no distinction
has been made notationally between random variables and particular observations of
those random variables: it is usually clear from the context which is meant.

1.3.1 Least squares estimation of β

Point estimates of the linear model parameters, β, can be obtained by the method of
least squares, that is by minimizing
n
X
S= (yi − µi )2 ,
i=1

with respect to β. To use least squares with a linear model, written in general matrix-
vector form, first recall the link between the Euclidean length of a vector and the sum
of squares of its elements. If v is any vector of dimension, n, then
n
X
2
kvk ≡ v v ≡T
vi2 .
i=1

Hence
S = ky − µk2 = ky − Xβk2

Since this expression is simply the squared (Euclidian) length of the vector y − Xβ,
its value will be unchanged if y − Xβ is rotated. This observation is the basis for a
THE THEORY OF LINEAR MODELS 13
practical method for finding β̂, and for developing the distributional results required
to use linear models.
Specifically, like any real matrix, X can always be decomposed
 
R
X=Q = Qf R (1.5)
0
where R is a p × p upper triangular matrix† , and Q is an n × n orthogonal matrix,
the first p columns of which form Qf (see A.6). Recall that orthogonal matrices
rotate vectors, but do not change their length. Orthogonality also means that QQT =
QT Q = In . Applying QT to y − Xβ implies that
  2
R
ky − Xβk2 = kQT y − QT Xβk2 = QT y − β .
0
 
f
Writing QT y = , where f is vector of dimension p, and hence r is a vector of
r
dimension n − p, yields
    2
2 f R
ky − Xβk = − β = kf − Rβk2 + krk2 .‡
r 0
The length of r does not depend on β, while kf − Rβk2 can be reduced to zero by
choosing β so that Rβ equals f . Hence
β̂ = R−1 f (1.6)
2 2
is the least squares estimator of β. Notice that krk = ky − Xβ̂k , the residual sum
of squares for the model fit.

1.3.2 The distribution of β̂

The distribution of the estimator, β̂, follows from that of QT y. Multivariate normal-
ity of QT y follows from that of y, and since the covariance matrix of y is In σ 2 , the
covariance matrix of QT y is
VQT y = QT In Qσ 2 = In σ 2 .
Furthermore    
f R
E = E(QT y) = QT Xβ = β
r 0
⇒ E(f ) = Rβ and E(r) = 0.
i.e. we have that
f ∼ N (Rβ, Ip σ 2 ) and r ∼ N (0, In−p σ 2 )
† i.e. R = 0 if i > j.
i,j » –
‡ If the last equality isn’t obvious recall that kxk2 = P x2 , so if x = v
, kxk2 =
P 2
i i w i vi +
P 2 2 2
i wi = kvk + kwk .
14 LINEAR MODELS
with both vectors independent of each other.
Turning to the properties of β̂ itself, unbiasedness follows immediately:
E(β̂) = R−1 E(f ) = R−1 Rβ = β.
Since the covariance matrix of f is Ip σ 2 , it also follows that the covariance matrix of
β̂ is
Vβ̂ = R−1 Ip R−T σ 2 = R−1 R−T σ 2 . (1.7)
Furthermore, since β̂ is just a linear transformation of the normal random variables
f , it must follow a multivariate normal distribution,
β̂ ∼ N (β, Vβ̂ ).
The forgoing distributional result is not usually directly useful for making inferences
about β, since σ 2 is generally unknown and must be estimated, thereby introducing
an extra component of variability that should be accounted for.

1.3.3 (β̂i − βi )/σ̂β̂i ∼ tn−p

Since the elements of r are i.i.d. N (0, σ 2 ) random variables,


n−p
1 2 1 X 2
krk = r ∼ χ2n−p .
σ2 σ 2 i=1 i

The mean of a χ2n−p r.v. is n − p, so this result is sufficient (but not necessary: see
exercise 7) to imply that
σ̂ 2 = krk2 /(n − p) (1.8)
is an unbiased estimator of σ 2§ . The independence of the elements of r and f also
implies that β̂ and σ̂ 2 are independent.
Now consider a single parameter estimator, β̂i , with standard deviation, σβ̂i , given
by the square root of element i, i of Vβ̂ . An unbiased estimator of Vβ̂ is V̂β̂ =
Vβ̂ σ̂ 2 /σ 2 = R−1 R−T σ̂ 2 , so an estimator, σ̂β̂i , is given by the square root of element
i, i of V̂β̂ , and it is clear that σ̂β̂i = σβ̂i σ̂/σ. Hence

β̂i − βi β̂i − βi (β̂i − βi )/σβ̂i N (0, 1)


= =q ∼q ∼ tn−p
σ̂β̂i σβ̂i σ̂/σ 1
krk2 /(n − p) χ 2 /(n − p)
σ 2 n−p

(where the independence of β̂i and σ̂ 2 has been used). This result enables confidence
intervals for βi to be found, and is the basis for hypothesis tests about individual βi ’s
(for example H0 : βi = 0).

§ Don’t forget that krk2 = ky − Xβ̂k2 .


THE THEORY OF LINEAR MODELS 15
1.3.4 F-ratio results

It is also of interest to obtain distributional results for testing, for example, the si-
multaneous equality to zero of several model parameters. Such tests are particularly
useful for making inferences about factor variables and their interactions, since each
factor (or interaction) is typically represented by several elements of β.
First consider partitioning the model matrix into two parts so that X = [X0 : X1 ],
where X0 and X1 have p − q and q columns, respectively. Let β0 and β1 be the
corresponding sub-vectors of β.
 
R0
QT X0 =
0
where R0 is the first p − q rows and columns of R (Q and R are from (1.5)). So
rotating y − X0 β0 using QT implies that
ky − X0 β0 k2 = kf0 − R0 β0 k2 + kf1 k2 + krk2
 
f0
where f has been partitioned so that f = (f1 being of dimension q). Hence
f1
2
kf1 k is the increase in residual sum of squares that results from dropping X1 from
the model (i.e. setting β1 = 0).
Now, under
H0 : β1 = 0,
E(f1 ) = 0 (using the facts that E(f ) = Rβ and R is upper triangular), and we
already know that the elements of f1 are i.i.d. normal r.v.s with variance σ 2 . Hence,
if H0 is true,
1
kf1 k2 ∼ χ2q .
σ2
So, forming an ‘F-ratio statistic’, F , assuming H0 , and recalling the independence of
f and r we have
kf1 k2 /q 1 2
σ 2 kf1 k /q
χ2q /q
F = = 1 ∼ ∼ Fq,n−p
σ̂ 2 2
σ 2 krk /(n − p) χ2n−p /(n − p)
and this result can be used to test the significance of model terms.
In general if β is partitioned into sub-vectors β0 , β1 . . . , βm (each usually relating
to a different effect), of dimensions q0 , q1 , . . . , qm , then f can also be so partitioned,
f T = [f0T , f1T , . . . , fm
T
], and tests of
H0 : βj = 0 vs. H1 : βj 6= 0
are conducted using the result that under H0
kfj k2 /qj
F = ∼ Fqj ,n−p
σ̂ 2
with F larger than this suggests, if the alternative is true. This is the result used to
draw up sequential ANOVA tables for a fitted model, of the sort produced by a single
16 LINEAR MODELS
argument call to anova in R. Note, however, that the hypothesis test about βj is
only valid in general if βk = 0 for all k such that j < k ≤ m: this follows from
the way that the test was derived, and is the reason that the ANOVA tables resulting
from such procedures are referred to as ‘sequential’ tables. The practical upshot is
that, if models are reduced in a different order, the p-values obtained will be different.
The exception to this is if the β̂j ’s are mutually independent, in which case the all
tests are simultaneously valid, and the ANOVA table for a model is not dependent on
the order of model terms: such independent β̂j ’s usually arise only in the context of
balanced data, from designed experiments.
Notice that sequential ANOVA tables are very easy to calculate: once a model has
been fitted by the QR method, all the relevant ‘sums of squares’ are easily calculated
directly from the elements of f , with the elements of r providing the residual sum of
squares.

1.3.5 The influence matrix

One matrix which will feature extensively in the discussion of GAMs is the influence
matrix (or hat matrix) of a linear model. This is the matrix which yields the fitted
value vector, µ̂, when post-multiplied by the data vector, y. Recalling the definition
of Qf , as being the first p columns of Q, f = Qf y and so
β̂ = R−1 QT
f y.

Furthermore µ̂ = Xβ̂ and X = Qf R so


µ̂ = Qf RR−1 QT T
f y = Qf Qf y

i.e. the matrix A ≡ Qf QT


f is the influence (hat) matrix such that µ̂ = Ay.

The influence matrix has a couple of interesting properties. Firstly, the trace of the
influence matrix is the number of (identifiable) parameters in the model, since
tr (A) = tr Qf QT T
 
f = tr Qf Qf = tr (Ip ) = p.

Secondly, AA = A, a property known as idempotency . Again the proof is simple:


AA = Qf QT T T T
f Qf Qf = Qf Ip Qf = Qf Qf = A.

1.3.6 The residuals, ˆ, and fitted values, µ̂

The influence matrix is helpful in deriving properties of the fitted values, µ̂, and
residuals, ˆ. µ̂ is unbiased, since E(µ̂) = E(Xβ̂) = XE(β̂) = Xβ = µ. The
covariance matrix of the fitted values is obtained from the fact that µ̂ is a linear
transformation of the random vector y, which has covariance matrix In σ 2 , so that
Vµ̂ = AIn AT σ 2 = Aσ 2 ,
by the idempotence (and symmetry) of A. The distribution of µ̂ is degenerate multi-
variate normal.
THE THEORY OF LINEAR MODELS 17
Similar arguments apply to the residuals.
ˆ = (I − A)y,
so
E(ˆ
) = E(y) − E(µ̂) = µ − µ = 0.
As in the fitted value case, we have
Vˆ = (In − A)In (In − A)T σ 2 = (In − 2A + AA) σ 2 = (In − A) σ 2
Again, the distribution of the residuals will be degenerate normal. The results for the
residuals are useful for model checking, since they allow the residuals to be stan-
dardized, so that they should have constant variance, if the model is correct.

1.3.7 Results in terms of X

The presentation so far has been in terms of the method actually used to fit linear
models in practice (the QR decomposition approach¶ ), which also greatly facilitates
the derivation of the distributional results required for practical modelling. However,
for historical reasons, these results are more usually presented in terms of the model
matrix, X, rather than the components of its QR decomposition. For completeness
some of the results are restated here, in terms of X.
Firstly consider the covariance matrix of β. This turns out to be (XT X)−1 σ 2 , which
is easily seen to be equivalent to (1.7) as follows:
−1 2 −1 2
Vβ̂ = (XT X)−1 σ 2 = RT QT f Qf R σ = RT R σ = R−1 R−T σ 2 .

The expression for the least squares estimates is β̂ = (XT X)−1 XT y, which is equiv-
alent to (1.6):
β̂ = (XT X)−1 XT y = R−1 R−T RT QT −1 T −1
f y = R Qf y = R f .

Given this last result it is easy to see that the influence matrix can be written:
A = X(XT X)−1 XT .
These results are of largely historical and theoretical interest: they should not be used
for computational purposes, and derivation of the distributional results is much more
difficult if one starts from these formulae.

1.3.8 The Gauss Markov Theorem: what’s special about least squares?

How good are least squares estimators? In particular, might it be possible to find
better estimators, in the sense of having lower variance while still being unbiased?

¶ A few programs still fit models by solution of XT Xβ̂ = XT y, but this is less computationally stable
than the rotation method described here, although it is a bit faster.
18 LINEAR MODELS
The Gauss Markov theorem shows that least squares estimators have the lowest vari-
ance of all unbiased estimators that are linear (meaning that the data only enter the
estimation process in a linear way).
Theorem: Suppose that µ ≡ E(Y) = Xβ and Vy = σ 2 I, and let φ̃ = cT Y be any
unbiased linear estimator of φ = tT β, where t is an arbitrary vector. Then:
Var(φ̃) ≥ Var(φ̂)
where φ̂ = t β̂, and β̂ = (X X)−1 XT Y is the least squares estimator of β. Notice
T T

that, since t is arbitrary, this theorem implies that each element of β̂ is a minimum
variance unbiased estimator.
Proof: Since φ̃ is a linear transformation of Y, Var(φ̃) = cT cσ 2 . To compare vari-
ances of φ̂ and φ̃ it is also useful to express Var(φ̂) in terms of c. To do this, note
that unbiasedness of φ̃ implies that
E(cT Y) = tT β ⇒ cT E(Y) = tT β ⇒ cT Xβ = tT β ⇒ cT X = tT .
So the variance of φ̂ can be written as
Var(φ̂) = Var(tT β̂) = Var(cT Xβ̂).
This is the variance of a linear transformation of β̂, and the covariance matrix of β̂
is (XT X)−1 σ 2 , so
Var(φ̂) = Var(cT Xβ̂) = cT X(XT X)−1 XT cσ 2 = cT Acσ 2
(where A is the influence or hat matrix). Now the variances of the two estimators can
be directly compared, and it can be seen that
Var(φ̃) ≥ Var(φ̂)
iff
cT (I − A)c ≥ 0.
This condition will always be met, because it is equivalent to:
[(I − A)c]T (I − A)c ≥ 0
by the idempotency of (I − A), but this last condition is saying that a sum of squares
can not be less than 0, which is clearly true. 2
Notice that this theorem uses independence and equal variance assumptions, but does
not assume normality. Of course there is a sense in which the theorem is intuitively
rather unsurprising, since it says that the minimum variance estimators are those
obtained by seeking to minimize the residual variance.

1.4 The geometry of linear modelling

A full understanding of what is happening when models are fitted by least squares
is facilitated by taking a geometric view of the fitting process. Some of the results
derived in the last few sections become rather obvious when viewed in this way.
THE GEOMETRY OF LINEAR MODELLING 19

0.8
3

0.6
0.4
2

3
0.2

2
1
0.0

0.0 0.2 0.4 0.6 0.8 1.0


1
x

Figure 1.4 The geometry of least squares. The left panel shows a straight line model fitted to
3 data by least squares. The right panel gives a geometric interpretation of the fitting process.
The 3 dimensional space shown is spanned by 3 orthogonal axes: one for each response vari-
able. The observed response vector, y, is shown as a point (•) within this space. The columns
of the model matrix define two directions within the space: the thick and dashed lines from
the origin. The model states that E(y) could be any linear combination of these vectors: i.e.
anywhere in the ‘model subspace’ indicated by the grey plane. Least squares fitting finds the
closest point in the model sub space to the response data (•): the ‘fitted values’. The short
thick line joins the response data to the fitted values: it is the ‘residual vector’.

1.4.1 Least squares

Again consider the linear model,


µ = Xβ, y ∼ N (µ, In σ 2 ),
where X is an n × p model matrix. But now consider an n dimensional Euclidean
space, <n , in which y defines the location of a single point. The space of all possible
linear combinations of the columns of X defines a subspace of <n , the elements of
this space being given by Xβ, where β can take any value in <p : this space will be
referred to as the space of X (strictly the column space). So, a linear model states
that, µ, the expected value of Y, lies in the space of X. Estimating a linear model,
by least squares, amounts to finding the point, µ̂ ≡ Xβ̂, in the space of X, that is
closest to the observed data y. Equivalently, µ̂ is the orthogonal projection of y on
to the space of X. An obvious, but important, consequence of this is that the residual
vector, ˆ = y − µ̂, is orthogonal to all vectors in the space of X.
Figure 1.4 illustrates these ideas for the simple case of a straight line regression
model for 3 data (shown conventionally in the left hand panel). The response data
and model are
   
.04 1 0.2  
β1
y =  .41  and µ =  1 1.0  .
β2
.62 1 0.6
Since β is unknown, the model simply says that µ could be any linear combination of
20 LINEAR MODELS

3
2

2
1 1

Figure 1.5 The geometry of fitting via orthogonal decompositions. The left panel illustrates
the geometry of the simple straight line model of 3 data introduced in figure 1.4. The right
hand panel shows how this original problem appears after rotation by QT , the transpose of
the orthogonal factor in a QR decomposition of X. Notice that in the rotated problem the
model subspace only has non-zero components relative to p axes (2 axes for this example),
while the residual vector has only zero components relative to those same axes.

the vectors [1, 1, 1]T and [.2, 1, .6]T . As the right hand panel of figure 1.4 illustrates,
fitting the model by least squares amounts to finding the particular linear combination
of the columns of these vectors, that is as close to y as possible (in terms of Euclidean
distance).

1.4.2 Fitting by orthogonal decompositions

Recall that the actual calculation of least squares estimates involves first forming the
QR decomposition of the model matrix, so that
 
R
X=Q ,
0
where Q is an n × n orthogonal matrix and R is a p × p upper triangular matrix.
Orthogonal matrices rotate vectors (without changing their length) and the first step
in least squares estimation is to rotate both the response vector, y, and the columns
of the model matrix, X, in exactly the same way, by pre-multiplication with QTk .
Figure 1.5 illustrates this rotation for the example shown in figure 1.4. The left panel
shows the response data and model space, for the original problem, while the right

k In fact the QR decomposition is not uniquely defined, in that the sign of rows of Q, and corresponding
columns of R, can be switched, without changing X — these sign changes are equivalent to reflections
of vectors, and the sign leading to maximum numerical stability is usually selected in practice. These
reflections don’t introduce any extra conceptual difficulty, but can make plots less easy to understand,
so I have surpressed them in this example.
THE GEOMETRY OF LINEAR MODELLING 21

3
2

2
1 1

Figure 1.6 The geometry of nested models.

hand panel shows the data and space after rotation by QT . Notice that, since the
problem has simply been rotated, the relative position of the data and basis vectors
(columns of X) has not changed. What has changed is that the problem now has
a particularly convenient orientation relative to the axes. The first two components
of the fitted value vector can now be read directly from axes 1 and 2, while the
third component is simply zero. By contrast, the residual vector has zero components
relative to axes 1 and 2, and its non-zero component can be read directly from axis
3. In terms of section 1.3.1, these vectors are [f T , 0T ]T and [0T , rT ]T , respectively.
The β̂ corresponding to the fitted values is now easily obtained. Of course we usually
require fitted values and residuals to be expressed in terms of the un-rotated problem,
but this is simply a matter of reversing the rotation using Q. i.e.
   
f 0
µ̂ = Q , and ˆ = Q .
0 r

1.4.3 Comparison of nested models

A linear model with model matrix X0 is nested within a linear model with model
matrix X1 if they are models for the same response data, and the columns of X0
span a subspace of the space spanned by the columns of X1 . Usually this simply
means that X1 is X0 with some extra columns added.
The vector of the difference between the fitted values of two nested linear models
is entirely within the subspace of the larger model, and is therefore orthogonal to
the residual vector for the larger model. This fact is geometrically obvious, as figure
1.6 illustrates, but it is a key reason why F ratio statistics have a relatively simple
distribution (under the simpler model).
Figure 1.6 is again based on the same simple straight line model that forms the ba-
sis for figures 1.4 and 1.5, but this time also illustrates the least squares fit of the
22 LINEAR MODELS
simplified model
yi = β0 + i ,
which is nested within the original straight line model. Again, both the original and
rotated versions of the model and data are shown. This time the fine continuous line
shows the projection of the response data onto the space of the simpler model, while
the fine dashed line shows the vector of the difference in fitted values between the
two models. Notice how this vector is orthogonal both to the reduced model subspace
and the full model residual vector.
The right panel of figure 1.6 illustrates that the rotation, using the transpose of the
orthogonal factor Q, of the full model matrix, has also lined up the problem very
conveniently for estimation of the reduced model. The fitted value vector for the
reduced model now has only one non-zero component, which is the component of
the rotated response data (•) relative to axis 1. The residual vector has gained the
component that the fitted value vector has lost, so it has zero component relative to
axis one, while its other components are the positions of the rotated response data
relative to axes 2 and 3.
So, much of the work required for estimating the simplified model has already been
done, when estimating the full model. Note, however, that if our interest had been in
comparing the full model to the model
yi = β1 xi + i ,
then it would have been necessary to reorder the columns of the full model matrix,
in order to avoid extra work in this way.

1.5 Practical linear models

This section covers practical linear modelling, via an extended example: the analysis
of data reported by Baker and Bellis (1993), which they used to support a theory
of ‘sperm competition’ in humans. The basic idea is that it is evolutionarily advan-
tageous for males to (sub-conciously) increase their sperm count in proportion to
the opportunities that their mate may have had for infidelity. Such behaviour has
been demonstrated in a wide variety of other animals, and using a sample of student
and staff volunteers from Manchester University, Baker and Bellis set out to see if
there is evidence for similar behaviour in humans. Two sets of data will be examined:
sperm.comp1 contains data on sperm count, time since last copulation and propor-
tion of that time spent together, for single copulations, from 15 heterosexual couples;
sperm.comp2 contains data on median sperm count, over multiple copulations,
for 24 heterosexual couples, along with the weight, height and age of the male and
female of each couple, and the volume of one teste of the male. From these data,
Baker and Bellis concluded that sperm count increases with the proportion of time,
since last copulation, that a couple have spent apart, and that sperm count increases
with female weight.
In general, practical linear modelling is concerned with finding an appropriate model
Another random document with
no related content on Scribd:
DANCE ON STILTS AT THE GIRLS’ UNYAGO, NIUCHI

Newala, too, suffers from the distance of its water-supply—at least


the Newala of to-day does; there was once another Newala in a lovely
valley at the foot of the plateau. I visited it and found scarcely a trace
of houses, only a Christian cemetery, with the graves of several
missionaries and their converts, remaining as a monument of its
former glories. But the surroundings are wonderfully beautiful. A
thick grove of splendid mango-trees closes in the weather-worn
crosses and headstones; behind them, combining the useful and the
agreeable, is a whole plantation of lemon-trees covered with ripe
fruit; not the small African kind, but a much larger and also juicier
imported variety, which drops into the hands of the passing traveller,
without calling for any exertion on his part. Old Newala is now under
the jurisdiction of the native pastor, Daudi, at Chingulungulu, who,
as I am on very friendly terms with him, allows me, as a matter of
course, the use of this lemon-grove during my stay at Newala.
FEET MUTILATED BY THE RAVAGES OF THE “JIGGER”
(Sarcopsylla penetrans)

The water-supply of New Newala is in the bottom of the valley,


some 1,600 feet lower down. The way is not only long and fatiguing,
but the water, when we get it, is thoroughly bad. We are suffering not
only from this, but from the fact that the arrangements at Newala are
nothing short of luxurious. We have a separate kitchen—a hut built
against the boma palisade on the right of the baraza, the interior of
which is not visible from our usual position. Our two cooks were not
long in finding this out, and they consequently do—or rather neglect
to do—what they please. In any case they do not seem to be very
particular about the boiling of our drinking-water—at least I can
attribute to no other cause certain attacks of a dysenteric nature,
from which both Knudsen and I have suffered for some time. If a
man like Omari has to be left unwatched for a moment, he is capable
of anything. Besides this complaint, we are inconvenienced by the
state of our nails, which have become as hard as glass, and crack on
the slightest provocation, and I have the additional infliction of
pimples all over me. As if all this were not enough, we have also, for
the last week been waging war against the jigger, who has found his
Eldorado in the hot sand of the Makonde plateau. Our men are seen
all day long—whenever their chronic colds and the dysentery likewise
raging among them permit—occupied in removing this scourge of
Africa from their feet and trying to prevent the disastrous
consequences of its presence. It is quite common to see natives of
this place with one or two toes missing; many have lost all their toes,
or even the whole front part of the foot, so that a well-formed leg
ends in a shapeless stump. These ravages are caused by the female of
Sarcopsylla penetrans, which bores its way under the skin and there
develops an egg-sac the size of a pea. In all books on the subject, it is
stated that one’s attention is called to the presence of this parasite by
an intolerable itching. This agrees very well with my experience, so
far as the softer parts of the sole, the spaces between and under the
toes, and the side of the foot are concerned, but if the creature
penetrates through the harder parts of the heel or ball of the foot, it
may escape even the most careful search till it has reached maturity.
Then there is no time to be lost, if the horrible ulceration, of which
we see cases by the dozen every day, is to be prevented. It is much
easier, by the way, to discover the insect on the white skin of a
European than on that of a native, on which the dark speck scarcely
shows. The four or five jiggers which, in spite of the fact that I
constantly wore high laced boots, chose my feet to settle in, were
taken out for me by the all-accomplished Knudsen, after which I
thought it advisable to wash out the cavities with corrosive
sublimate. The natives have a different sort of disinfectant—they fill
the hole with scraped roots. In a tiny Makua village on the slope of
the plateau south of Newala, we saw an old woman who had filled all
the spaces under her toe-nails with powdered roots by way of
prophylactic treatment. What will be the result, if any, who can say?
The rest of the many trifling ills which trouble our existence are
really more comic than serious. In the absence of anything else to
smoke, Knudsen and I at last opened a box of cigars procured from
the Indian store-keeper at Lindi, and tried them, with the most
distressing results. Whether they contain opium or some other
narcotic, neither of us can say, but after the tenth puff we were both
“off,” three-quarters stupefied and unspeakably wretched. Slowly we
recovered—and what happened next? Half-an-hour later we were
once more smoking these poisonous concoctions—so insatiable is the
craving for tobacco in the tropics.
Even my present attacks of fever scarcely deserve to be taken
seriously. I have had no less than three here at Newala, all of which
have run their course in an incredibly short time. In the early
afternoon, I am busy with my old natives, asking questions and
making notes. The strong midday coffee has stimulated my spirits to
an extraordinary degree, the brain is active and vigorous, and work
progresses rapidly, while a pleasant warmth pervades the whole
body. Suddenly this gives place to a violent chill, forcing me to put on
my overcoat, though it is only half-past three and the afternoon sun
is at its hottest. Now the brain no longer works with such acuteness
and logical precision; more especially does it fail me in trying to
establish the syntax of the difficult Makua language on which I have
ventured, as if I had not enough to do without it. Under the
circumstances it seems advisable to take my temperature, and I do
so, to save trouble, without leaving my seat, and while going on with
my work. On examination, I find it to be 101·48°. My tutors are
abruptly dismissed and my bed set up in the baraza; a few minutes
later I am in it and treating myself internally with hot water and
lemon-juice.
Three hours later, the thermometer marks nearly 104°, and I make
them carry me back into the tent, bed and all, as I am now perspiring
heavily, and exposure to the cold wind just beginning to blow might
mean a fatal chill. I lie still for a little while, and then find, to my
great relief, that the temperature is not rising, but rather falling. This
is about 7.30 p.m. At 8 p.m. I find, to my unbounded astonishment,
that it has fallen below 98·6°, and I feel perfectly well. I read for an
hour or two, and could very well enjoy a smoke, if I had the
wherewithal—Indian cigars being out of the question.
Having no medical training, I am at a loss to account for this state
of things. It is impossible that these transitory attacks of high fever
should be malarial; it seems more probable that they are due to a
kind of sunstroke. On consulting my note-book, I become more and
more inclined to think this is the case, for these attacks regularly
follow extreme fatigue and long exposure to strong sunshine. They at
least have the advantage of being only short interruptions to my
work, as on the following morning I am always quite fresh and fit.
My treasure of a cook is suffering from an enormous hydrocele which
makes it difficult for him to get up, and Moritz is obliged to keep in
the dark on account of his inflamed eyes. Knudsen’s cook, a raw boy
from somewhere in the bush, knows still less of cooking than Omari;
consequently Nils Knudsen himself has been promoted to the vacant
post. Finding that we had come to the end of our supplies, he began
by sending to Chingulungulu for the four sucking-pigs which we had
bought from Matola and temporarily left in his charge; and when
they came up, neatly packed in a large crate, he callously slaughtered
the biggest of them. The first joint we were thoughtless enough to
entrust for roasting to Knudsen’s mshenzi cook, and it was
consequently uneatable; but we made the rest of the animal into a
jelly which we ate with great relish after weeks of underfeeding,
consuming incredible helpings of it at both midday and evening
meals. The only drawback is a certain want of variety in the tinned
vegetables. Dr. Jäger, to whom the Geographical Commission
entrusted the provisioning of the expeditions—mine as well as his
own—because he had more time on his hands than the rest of us,
seems to have laid in a huge stock of Teltow turnips,[46] an article of
food which is all very well for occasional use, but which quickly palls
when set before one every day; and we seem to have no other tins
left. There is no help for it—we must put up with the turnips; but I
am certain that, once I am home again, I shall not touch them for ten
years to come.
Amid all these minor evils, which, after all, go to make up the
genuine flavour of Africa, there is at least one cheering touch:
Knudsen has, with the dexterity of a skilled mechanic, repaired my 9
× 12 cm. camera, at least so far that I can use it with a little care.
How, in the absence of finger-nails, he was able to accomplish such a
ticklish piece of work, having no tool but a clumsy screw-driver for
taking to pieces and putting together again the complicated
mechanism of the instantaneous shutter, is still a mystery to me; but
he did it successfully. The loss of his finger-nails shows him in a light
contrasting curiously enough with the intelligence evinced by the
above operation; though, after all, it is scarcely surprising after his
ten years’ residence in the bush. One day, at Lindi, he had occasion
to wash a dog, which must have been in need of very thorough
cleansing, for the bottle handed to our friend for the purpose had an
extremely strong smell. Having performed his task in the most
conscientious manner, he perceived with some surprise that the dog
did not appear much the better for it, and was further surprised by
finding his own nails ulcerating away in the course of the next few
days. “How was I to know that carbolic acid has to be diluted?” he
mutters indignantly, from time to time, with a troubled gaze at his
mutilated finger-tips.
Since we came to Newala we have been making excursions in all
directions through the surrounding country, in accordance with old
habit, and also because the akida Sefu did not get together the tribal
elders from whom I wanted information so speedily as he had
promised. There is, however, no harm done, as, even if seen only
from the outside, the country and people are interesting enough.
The Makonde plateau is like a large rectangular table rounded off
at the corners. Measured from the Indian Ocean to Newala, it is
about seventy-five miles long, and between the Rovuma and the
Lukuledi it averages fifty miles in breadth, so that its superficial area
is about two-thirds of that of the kingdom of Saxony. The surface,
however, is not level, but uniformly inclined from its south-western
edge to the ocean. From the upper edge, on which Newala lies, the
eye ranges for many miles east and north-east, without encountering
any obstacle, over the Makonde bush. It is a green sea, from which
here and there thick clouds of smoke rise, to show that it, too, is
inhabited by men who carry on their tillage like so many other
primitive peoples, by cutting down and burning the bush, and
manuring with the ashes. Even in the radiant light of a tropical day
such a fire is a grand sight.
Much less effective is the impression produced just now by the
great western plain as seen from the edge of the plateau. As often as
time permits, I stroll along this edge, sometimes in one direction,
sometimes in another, in the hope of finding the air clear enough to
let me enjoy the view; but I have always been disappointed.
Wherever one looks, clouds of smoke rise from the burning bush,
and the air is full of smoke and vapour. It is a pity, for under more
favourable circumstances the panorama of the whole country up to
the distant Majeje hills must be truly magnificent. It is of little use
taking photographs now, and an outline sketch gives a very poor idea
of the scenery. In one of these excursions I went out of my way to
make a personal attempt on the Makonde bush. The present edge of
the plateau is the result of a far-reaching process of destruction
through erosion and denudation. The Makonde strata are
everywhere cut into by ravines, which, though short, are hundreds of
yards in depth. In consequence of the loose stratification of these
beds, not only are the walls of these ravines nearly vertical, but their
upper end is closed by an equally steep escarpment, so that the
western edge of the Makonde plateau is hemmed in by a series of
deep, basin-like valleys. In order to get from one side of such a ravine
to the other, I cut my way through the bush with a dozen of my men.
It was a very open part, with more grass than scrub, but even so the
short stretch of less than two hundred yards was very hard work; at
the end of it the men’s calicoes were in rags and they themselves
bleeding from hundreds of scratches, while even our strong khaki
suits had not escaped scatheless.

NATIVE PATH THROUGH THE MAKONDE BUSH, NEAR


MAHUTA

I see increasing reason to believe that the view formed some time
back as to the origin of the Makonde bush is the correct one. I have
no doubt that it is not a natural product, but the result of human
occupation. Those parts of the high country where man—as a very
slight amount of practice enables the eye to perceive at once—has not
yet penetrated with axe and hoe, are still occupied by a splendid
timber forest quite able to sustain a comparison with our mixed
forests in Germany. But wherever man has once built his hut or tilled
his field, this horrible bush springs up. Every phase of this process
may be seen in the course of a couple of hours’ walk along the main
road. From the bush to right or left, one hears the sound of the axe—
not from one spot only, but from several directions at once. A few
steps further on, we can see what is taking place. The brush has been
cut down and piled up in heaps to the height of a yard or more,
between which the trunks of the large trees stand up like the last
pillars of a magnificent ruined building. These, too, present a
melancholy spectacle: the destructive Makonde have ringed them—
cut a broad strip of bark all round to ensure their dying off—and also
piled up pyramids of brush round them. Father and son, mother and
son-in-law, are chopping away perseveringly in the background—too
busy, almost, to look round at the white stranger, who usually excites
so much interest. If you pass by the same place a week later, the piles
of brushwood have disappeared and a thick layer of ashes has taken
the place of the green forest. The large trees stretch their
smouldering trunks and branches in dumb accusation to heaven—if
they have not already fallen and been more or less reduced to ashes,
perhaps only showing as a white stripe on the dark ground.
This work of destruction is carried out by the Makonde alike on the
virgin forest and on the bush which has sprung up on sites already
cultivated and deserted. In the second case they are saved the trouble
of burning the large trees, these being entirely absent in the
secondary bush.
After burning this piece of forest ground and loosening it with the
hoe, the native sows his corn and plants his vegetables. All over the
country, he goes in for bed-culture, which requires, and, in fact,
receives, the most careful attention. Weeds are nowhere tolerated in
the south of German East Africa. The crops may fail on the plains,
where droughts are frequent, but never on the plateau with its
abundant rains and heavy dews. Its fortunate inhabitants even have
the satisfaction of seeing the proud Wayao and Wamakua working
for them as labourers, driven by hunger to serve where they were
accustomed to rule.
But the light, sandy soil is soon exhausted, and would yield no
harvest the second year if cultivated twice running. This fact has
been familiar to the native for ages; consequently he provides in
time, and, while his crop is growing, prepares the next plot with axe
and firebrand. Next year he plants this with his various crops and
lets the first piece lie fallow. For a short time it remains waste and
desolate; then nature steps in to repair the destruction wrought by
man; a thousand new growths spring out of the exhausted soil, and
even the old stumps put forth fresh shoots. Next year the new growth
is up to one’s knees, and in a few years more it is that terrible,
impenetrable bush, which maintains its position till the black
occupier of the land has made the round of all the available sites and
come back to his starting point.
The Makonde are, body and soul, so to speak, one with this bush.
According to my Yao informants, indeed, their name means nothing
else but “bush people.” Their own tradition says that they have been
settled up here for a very long time, but to my surprise they laid great
stress on an original immigration. Their old homes were in the
south-east, near Mikindani and the mouth of the Rovuma, whence
their peaceful forefathers were driven by the continual raids of the
Sakalavas from Madagascar and the warlike Shirazis[47] of the coast,
to take refuge on the almost inaccessible plateau. I have studied
African ethnology for twenty years, but the fact that changes of
population in this apparently quiet and peaceable corner of the earth
could have been occasioned by outside enterprises taking place on
the high seas, was completely new to me. It is, no doubt, however,
correct.
The charming tribal legend of the Makonde—besides informing us
of other interesting matters—explains why they have to live in the
thickest of the bush and a long way from the edge of the plateau,
instead of making their permanent homes beside the purling brooks
and springs of the low country.
“The place where the tribe originated is Mahuta, on the southern
side of the plateau towards the Rovuma, where of old time there was
nothing but thick bush. Out of this bush came a man who never
washed himself or shaved his head, and who ate and drank but little.
He went out and made a human figure from the wood of a tree
growing in the open country, which he took home to his abode in the
bush and there set it upright. In the night this image came to life and
was a woman. The man and woman went down together to the
Rovuma to wash themselves. Here the woman gave birth to a still-
born child. They left that place and passed over the high land into the
valley of the Mbemkuru, where the woman had another child, which
was also born dead. Then they returned to the high bush country of
Mahuta, where the third child was born, which lived and grew up. In
course of time, the couple had many more children, and called
themselves Wamatanda. These were the ancestral stock of the
Makonde, also called Wamakonde,[48] i.e., aborigines. Their
forefather, the man from the bush, gave his children the command to
bury their dead upright, in memory of the mother of their race who
was cut out of wood and awoke to life when standing upright. He also
warned them against settling in the valleys and near large streams,
for sickness and death dwelt there. They were to make it a rule to
have their huts at least an hour’s walk from the nearest watering-
place; then their children would thrive and escape illness.”
The explanation of the name Makonde given by my informants is
somewhat different from that contained in the above legend, which I
extract from a little book (small, but packed with information), by
Pater Adams, entitled Lindi und sein Hinterland. Otherwise, my
results agree exactly with the statements of the legend. Washing?
Hapana—there is no such thing. Why should they do so? As it is, the
supply of water scarcely suffices for cooking and drinking; other
people do not wash, so why should the Makonde distinguish himself
by such needless eccentricity? As for shaving the head, the short,
woolly crop scarcely needs it,[49] so the second ancestral precept is
likewise easy enough to follow. Beyond this, however, there is
nothing ridiculous in the ancestor’s advice. I have obtained from
various local artists a fairly large number of figures carved in wood,
ranging from fifteen to twenty-three inches in height, and
representing women belonging to the great group of the Mavia,
Makonde, and Matambwe tribes. The carving is remarkably well
done and renders the female type with great accuracy, especially the
keloid ornamentation, to be described later on. As to the object and
meaning of their works the sculptors either could or (more probably)
would tell me nothing, and I was forced to content myself with the
scanty information vouchsafed by one man, who said that the figures
were merely intended to represent the nembo—the artificial
deformations of pelele, ear-discs, and keloids. The legend recorded
by Pater Adams places these figures in a new light. They must surely
be more than mere dolls; and we may even venture to assume that
they are—though the majority of present-day Makonde are probably
unaware of the fact—representations of the tribal ancestress.
The references in the legend to the descent from Mahuta to the
Rovuma, and to a journey across the highlands into the Mbekuru
valley, undoubtedly indicate the previous history of the tribe, the
travels of the ancestral pair typifying the migrations of their
descendants. The descent to the neighbouring Rovuma valley, with
its extraordinary fertility and great abundance of game, is intelligible
at a glance—but the crossing of the Lukuledi depression, the ascent
to the Rondo Plateau and the descent to the Mbemkuru, also lie
within the bounds of probability, for all these districts have exactly
the same character as the extreme south. Now, however, comes a
point of especial interest for our bacteriological age. The primitive
Makonde did not enjoy their lives in the marshy river-valleys.
Disease raged among them, and many died. It was only after they
had returned to their original home near Mahuta, that the health
conditions of these people improved. We are very apt to think of the
African as a stupid person whose ignorance of nature is only equalled
by his fear of it, and who looks on all mishaps as caused by evil
spirits and malignant natural powers. It is much more correct to
assume in this case that the people very early learnt to distinguish
districts infested with malaria from those where it is absent.
This knowledge is crystallized in the
ancestral warning against settling in the
valleys and near the great waters, the
dwelling-places of disease and death. At the
same time, for security against the hostile
Mavia south of the Rovuma, it was enacted
that every settlement must be not less than a
certain distance from the southern edge of the
plateau. Such in fact is their mode of life at the
present day. It is not such a bad one, and
certainly they are both safer and more
comfortable than the Makua, the recent
intruders from the south, who have made USUAL METHOD OF
good their footing on the western edge of the CLOSING HUT-DOOR
plateau, extending over a fairly wide belt of
country. Neither Makua nor Makonde show in their dwellings
anything of the size and comeliness of the Yao houses in the plain,
especially at Masasi, Chingulungulu and Zuza’s. Jumbe Chauro, a
Makonde hamlet not far from Newala, on the road to Mahuta, is the
most important settlement of the tribe I have yet seen, and has fairly
spacious huts. But how slovenly is their construction compared with
the palatial residences of the elephant-hunters living in the plain.
The roofs are still more untidy than in the general run of huts during
the dry season, the walls show here and there the scanty beginnings
or the lamentable remains of the mud plastering, and the interior is a
veritable dog-kennel; dirt, dust and disorder everywhere. A few huts
only show any attempt at division into rooms, and this consists
merely of very roughly-made bamboo partitions. In one point alone
have I noticed any indication of progress—in the method of fastening
the door. Houses all over the south are secured in a simple but
ingenious manner. The door consists of a set of stout pieces of wood
or bamboo, tied with bark-string to two cross-pieces, and moving in
two grooves round one of the door-posts, so as to open inwards. If
the owner wishes to leave home, he takes two logs as thick as a man’s
upper arm and about a yard long. One of these is placed obliquely
against the middle of the door from the inside, so as to form an angle
of from 60° to 75° with the ground. He then places the second piece
horizontally across the first, pressing it downward with all his might.
It is kept in place by two strong posts planted in the ground a few
inches inside the door. This fastening is absolutely safe, but of course
cannot be applied to both doors at once, otherwise how could the
owner leave or enter his house? I have not yet succeeded in finding
out how the back door is fastened.

MAKONDE LOCK AND KEY AT JUMBE CHAURO


This is the general way of closing a house. The Makonde at Jumbe
Chauro, however, have a much more complicated, solid and original
one. Here, too, the door is as already described, except that there is
only one post on the inside, standing by itself about six inches from
one side of the doorway. Opposite this post is a hole in the wall just
large enough to admit a man’s arm. The door is closed inside by a
large wooden bolt passing through a hole in this post and pressing
with its free end against the door. The other end has three holes into
which fit three pegs running in vertical grooves inside the post. The
door is opened with a wooden key about a foot long, somewhat
curved and sloped off at the butt; the other end has three pegs
corresponding to the holes, in the bolt, so that, when it is thrust
through the hole in the wall and inserted into the rectangular
opening in the post, the pegs can be lifted and the bolt drawn out.[50]

MODE OF INSERTING THE KEY

With no small pride first one householder and then a second


showed me on the spot the action of this greatest invention of the
Makonde Highlands. To both with an admiring exclamation of
“Vizuri sana!” (“Very fine!”). I expressed the wish to take back these
marvels with me to Ulaya, to show the Wazungu what clever fellows
the Makonde are. Scarcely five minutes after my return to camp at
Newala, the two men came up sweating under the weight of two
heavy logs which they laid down at my feet, handing over at the same
time the keys of the fallen fortress. Arguing, logically enough, that if
the key was wanted, the lock would be wanted with it, they had taken
their axes and chopped down the posts—as it never occurred to them
to dig them out of the ground and so bring them intact. Thus I have
two badly damaged specimens, and the owners, instead of praise,
come in for a blowing-up.
The Makua huts in the environs of Newala are especially
miserable; their more than slovenly construction reminds one of the
temporary erections of the Makua at Hatia’s, though the people here
have not been concerned in a war. It must therefore be due to
congenital idleness, or else to the absence of a powerful chief. Even
the baraza at Mlipa’s, a short hour’s walk south-east of Newala,
shares in this general neglect. While public buildings in this country
are usually looked after more or less carefully, this is in evident
danger of being blown over by the first strong easterly gale. The only
attractive object in this whole district is the grave of the late chief
Mlipa. I visited it in the morning, while the sun was still trying with
partial success to break through the rolling mists, and the circular
grove of tall euphorbias, which, with a broken pot, is all that marks
the old king’s resting-place, impressed one with a touch of pathos.
Even my very materially-minded carriers seemed to feel something
of the sort, for instead of their usual ribald songs, they chanted
solemnly, as we marched on through the dense green of the Makonde
bush:—
“We shall arrive with the great master; we stand in a row and have
no fear about getting our food and our money from the Serkali (the
Government). We are not afraid; we are going along with the great
master, the lion; we are going down to the coast and back.”
With regard to the characteristic features of the various tribes here
on the western edge of the plateau, I can arrive at no other
conclusion than the one already come to in the plain, viz., that it is
impossible for anyone but a trained anthropologist to assign any
given individual at once to his proper tribe. In fact, I think that even
an anthropological specialist, after the most careful examination,
might find it a difficult task to decide. The whole congeries of peoples
collected in the region bounded on the west by the great Central
African rift, Tanganyika and Nyasa, and on the east by the Indian
Ocean, are closely related to each other—some of their languages are
only distinguished from one another as dialects of the same speech,
and no doubt all the tribes present the same shape of skull and
structure of skeleton. Thus, surely, there can be no very striking
differences in outward appearance.
Even did such exist, I should have no time
to concern myself with them, for day after day,
I have to see or hear, as the case may be—in
any case to grasp and record—an
extraordinary number of ethnographic
phenomena. I am almost disposed to think it
fortunate that some departments of inquiry, at
least, are barred by external circumstances.
Chief among these is the subject of iron-
working. We are apt to think of Africa as a
country where iron ore is everywhere, so to
speak, to be picked up by the roadside, and
where it would be quite surprising if the
inhabitants had not learnt to smelt the
material ready to their hand. In fact, the
knowledge of this art ranges all over the
continent, from the Kabyles in the north to the
Kafirs in the south. Here between the Rovuma
and the Lukuledi the conditions are not so
favourable. According to the statements of the
Makonde, neither ironstone nor any other
form of iron ore is known to them. They have
not therefore advanced to the art of smelting
the metal, but have hitherto bought all their
THE ANCESTRESS OF
THE MAKONDE
iron implements from neighbouring tribes.
Even in the plain the inhabitants are not much
better off. Only one man now living is said to
understand the art of smelting iron. This old fundi lives close to
Huwe, that isolated, steep-sided block of granite which rises out of
the green solitude between Masasi and Chingulungulu, and whose
jagged and splintered top meets the traveller’s eye everywhere. While
still at Masasi I wished to see this man at work, but was told that,
frightened by the rising, he had retired across the Rovuma, though
he would soon return. All subsequent inquiries as to whether the
fundi had come back met with the genuine African answer, “Bado”
(“Not yet”).
BRAZIER

Some consolation was afforded me by a brassfounder, whom I


came across in the bush near Akundonde’s. This man is the favourite
of women, and therefore no doubt of the gods; he welds the glittering
brass rods purchased at the coast into those massive, heavy rings
which, on the wrists and ankles of the local fair ones, continually give
me fresh food for admiration. Like every decent master-craftsman he
had all his tools with him, consisting of a pair of bellows, three
crucibles and a hammer—nothing more, apparently. He was quite
willing to show his skill, and in a twinkling had fixed his bellows on
the ground. They are simply two goat-skins, taken off whole, the four
legs being closed by knots, while the upper opening, intended to
admit the air, is kept stretched by two pieces of wood. At the lower
end of the skin a smaller opening is left into which a wooden tube is
stuck. The fundi has quickly borrowed a heap of wood-embers from
the nearest hut; he then fixes the free ends of the two tubes into an
earthen pipe, and clamps them to the ground by means of a bent
piece of wood. Now he fills one of his small clay crucibles, the dross
on which shows that they have been long in use, with the yellow
material, places it in the midst of the embers, which, at present are
only faintly glimmering, and begins his work. In quick alternation
the smith’s two hands move up and down with the open ends of the
bellows; as he raises his hand he holds the slit wide open, so as to let
the air enter the skin bag unhindered. In pressing it down he closes
the bag, and the air puffs through the bamboo tube and clay pipe into
the fire, which quickly burns up. The smith, however, does not keep
on with this work, but beckons to another man, who relieves him at
the bellows, while he takes some more tools out of a large skin pouch
carried on his back. I look on in wonder as, with a smooth round
stick about the thickness of a finger, he bores a few vertical holes into
the clean sand of the soil. This should not be difficult, yet the man
seems to be taking great pains over it. Then he fastens down to the
ground, with a couple of wooden clamps, a neat little trough made by
splitting a joint of bamboo in half, so that the ends are closed by the
two knots. At last the yellow metal has attained the right consistency,
and the fundi lifts the crucible from the fire by means of two sticks
split at the end to serve as tongs. A short swift turn to the left—a
tilting of the crucible—and the molten brass, hissing and giving forth
clouds of smoke, flows first into the bamboo mould and then into the
holes in the ground.
The technique of this backwoods craftsman may not be very far
advanced, but it cannot be denied that he knows how to obtain an
adequate result by the simplest means. The ladies of highest rank in
this country—that is to say, those who can afford it, wear two kinds
of these massive brass rings, one cylindrical, the other semicircular
in section. The latter are cast in the most ingenious way in the
bamboo mould, the former in the circular hole in the sand. It is quite
a simple matter for the fundi to fit these bars to the limbs of his fair
customers; with a few light strokes of his hammer he bends the
pliable brass round arm or ankle without further inconvenience to
the wearer.
SHAPING THE POT

SMOOTHING WITH MAIZE-COB

CUTTING THE EDGE


FINISHING THE BOTTOM

LAST SMOOTHING BEFORE


BURNING

FIRING THE BRUSH-PILE


LIGHTING THE FARTHER SIDE OF
THE PILE

TURNING THE RED-HOT VESSEL

NYASA WOMAN MAKING POTS AT MASASI


Pottery is an art which must always and everywhere excite the
interest of the student, just because it is so intimately connected with
the development of human culture, and because its relics are one of
the principal factors in the reconstruction of our own condition in
prehistoric times. I shall always remember with pleasure the two or
three afternoons at Masasi when Salim Matola’s mother, a slightly-
built, graceful, pleasant-looking woman, explained to me with
touching patience, by means of concrete illustrations, the ceramic art
of her people. The only implements for this primitive process were a
lump of clay in her left hand, and in the right a calabash containing
the following valuables: the fragment of a maize-cob stripped of all
its grains, a smooth, oval pebble, about the size of a pigeon’s egg, a
few chips of gourd-shell, a bamboo splinter about the length of one’s
hand, a small shell, and a bunch of some herb resembling spinach.
Nothing more. The woman scraped with the
shell a round, shallow hole in the soft, fine
sand of the soil, and, when an active young
girl had filled the calabash with water for her,
she began to knead the clay. As if by magic it
gradually assumed the shape of a rough but
already well-shaped vessel, which only wanted
a little touching up with the instruments
before mentioned. I looked out with the
MAKUA WOMAN closest attention for any indication of the use
MAKING A POT. of the potter’s wheel, in however rudimentary
SHOWS THE a form, but no—hapana (there is none). The
BEGINNINGS OF THE embryo pot stood firmly in its little
POTTER’S WHEEL
depression, and the woman walked round it in
a stooping posture, whether she was removing
small stones or similar foreign bodies with the maize-cob, smoothing
the inner or outer surface with the splinter of bamboo, or later, after
letting it dry for a day, pricking in the ornamentation with a pointed
bit of gourd-shell, or working out the bottom, or cutting the edge
with a sharp bamboo knife, or giving the last touches to the finished
vessel. This occupation of the women is infinitely toilsome, but it is
without doubt an accurate reproduction of the process in use among
our ancestors of the Neolithic and Bronze ages.
There is no doubt that the invention of pottery, an item in human
progress whose importance cannot be over-estimated, is due to
women. Rough, coarse and unfeeling, the men of the horde range
over the countryside. When the united cunning of the hunters has
succeeded in killing the game; not one of them thinks of carrying
home the spoil. A bright fire, kindled by a vigorous wielding of the
drill, is crackling beside them; the animal has been cleaned and cut
up secundum artem, and, after a slight singeing, will soon disappear
under their sharp teeth; no one all this time giving a single thought
to wife or child.
To what shifts, on the other hand, the primitive wife, and still more
the primitive mother, was put! Not even prehistoric stomachs could
endure an unvarying diet of raw food. Something or other suggested
the beneficial effect of hot water on the majority of approved but
indigestible dishes. Perhaps a neighbour had tried holding the hard
roots or tubers over the fire in a calabash filled with water—or maybe
an ostrich-egg-shell, or a hastily improvised vessel of bark. They
became much softer and more palatable than they had previously
been; but, unfortunately, the vessel could not stand the fire and got
charred on the outside. That can be remedied, thought our
ancestress, and plastered a layer of wet clay round a similar vessel.
This is an improvement; the cooking utensil remains uninjured, but
the heat of the fire has shrunk it, so that it is loose in its shell. The
next step is to detach it, so, with a firm grip and a jerk, shell and
kernel are separated, and pottery is invented. Perhaps, however, the
discovery which led to an intelligent use of the burnt-clay shell, was
made in a slightly different way. Ostrich-eggs and calabashes are not
to be found in every part of the world, but everywhere mankind has
arrived at the art of making baskets out of pliant materials, such as
bark, bast, strips of palm-leaf, supple twigs, etc. Our inventor has no
water-tight vessel provided by nature. “Never mind, let us line the
basket with clay.” This answers the purpose, but alas! the basket gets
burnt over the blazing fire, the woman watches the process of
cooking with increasing uneasiness, fearing a leak, but no leak
appears. The food, done to a turn, is eaten with peculiar relish; and
the cooking-vessel is examined, half in curiosity, half in satisfaction
at the result. The plastic clay is now hard as stone, and at the same
time looks exceedingly well, for the neat plaiting of the burnt basket
is traced all over it in a pretty pattern. Thus, simultaneously with
pottery, its ornamentation was invented.
Primitive woman has another claim to respect. It was the man,
roving abroad, who invented the art of producing fire at will, but the
woman, unable to imitate him in this, has been a Vestal from the
earliest times. Nothing gives so much trouble as the keeping alight of
the smouldering brand, and, above all, when all the men are absent
from the camp. Heavy rain-clouds gather, already the first large
drops are falling, the first gusts of the storm rage over the plain. The
little flame, a greater anxiety to the woman than her own children,
flickers unsteadily in the blast. What is to be done? A sudden thought
occurs to her, and in an instant she has constructed a primitive hut
out of strips of bark, to protect the flame against rain and wind.
This, or something very like it, was the way in which the principle
of the house was discovered; and even the most hardened misogynist
cannot fairly refuse a woman the credit of it. The protection of the
hearth-fire from the weather is the germ from which the human
dwelling was evolved. Men had little, if any share, in this forward
step, and that only at a late stage. Even at the present day, the
plastering of the housewall with clay and the manufacture of pottery
are exclusively the women’s business. These are two very significant
survivals. Our European kitchen-garden, too, is originally a woman’s
invention, and the hoe, the primitive instrument of agriculture, is,
characteristically enough, still used in this department. But the
noblest achievement which we owe to the other sex is unquestionably
the art of cookery. Roasting alone—the oldest process—is one for
which men took the hint (a very obvious one) from nature. It must
have been suggested by the scorched carcase of some animal
overtaken by the destructive forest-fires. But boiling—the process of
improving organic substances by the help of water heated to boiling-
point—is a much later discovery. It is so recent that it has not even
yet penetrated to all parts of the world. The Polynesians understand
how to steam food, that is, to cook it, neatly wrapped in leaves, in a
hole in the earth between hot stones, the air being excluded, and
(sometimes) a few drops of water sprinkled on the stones; but they
do not understand boiling.
To come back from this digression, we find that the slender Nyasa
woman has, after once more carefully examining the finished pot,
put it aside in the shade to dry. On the following day she sends me
word by her son, Salim Matola, who is always on hand, that she is
going to do the burning, and, on coming out of my house, I find her
already hard at work. She has spread on the ground a layer of very
dry sticks, about as thick as one’s thumb, has laid the pot (now of a
yellowish-grey colour) on them, and is piling brushwood round it.
My faithful Pesa mbili, the mnyampara, who has been standing by,
most obligingly, with a lighted stick, now hands it to her. Both of
them, blowing steadily, light the pile on the lee side, and, when the
flame begins to catch, on the weather side also. Soon the whole is in a
blaze, but the dry fuel is quickly consumed and the fire dies down, so
that we see the red-hot vessel rising from the ashes. The woman
turns it continually with a long stick, sometimes one way and
sometimes another, so that it may be evenly heated all over. In
twenty minutes she rolls it out of the ash-heap, takes up the bundle
of spinach, which has been lying for two days in a jar of water, and
sprinkles the red-hot clay with it. The places where the drops fall are
marked by black spots on the uniform reddish-brown surface. With a
sigh of relief, and with visible satisfaction, the woman rises to an
erect position; she is standing just in a line between me and the fire,
from which a cloud of smoke is just rising: I press the ball of my
camera, the shutter clicks—the apotheosis is achieved! Like a
priestess, representative of her inventive sex, the graceful woman
stands: at her feet the hearth-fire she has given us beside her the
invention she has devised for us, in the background the home she has
built for us.
At Newala, also, I have had the manufacture of pottery carried on
in my presence. Technically the process is better than that already
described, for here we find the beginnings of the potter’s wheel,
which does not seem to exist in the plains; at least I have seen
nothing of the sort. The artist, a frightfully stupid Makua woman, did
not make a depression in the ground to receive the pot she was about
to shape, but used instead a large potsherd. Otherwise, she went to
work in much the same way as Salim’s mother, except that she saved
herself the trouble of walking round and round her work by squatting
at her ease and letting the pot and potsherd rotate round her; this is
surely the first step towards a machine. But it does not follow that
the pot was improved by the process. It is true that it was beautifully
rounded and presented a very creditable appearance when finished,
but the numerous large and small vessels which I have seen, and, in
part, collected, in the “less advanced” districts, are no less so. We
moderns imagine that instruments of precision are necessary to
produce excellent results. Go to the prehistoric collections of our
museums and look at the pots, urns and bowls of our ancestors in the
dim ages of the past, and you will at once perceive your error.
MAKING LONGITUDINAL CUT IN
BARK

DRAWING THE BARK OFF THE LOG

REMOVING THE OUTER BARK


BEATING THE BARK

WORKING THE BARK-CLOTH AFTER BEATING, TO MAKE IT


SOFT

MANUFACTURE OF BARK-CLOTH AT NEWALA


To-day, nearly the whole population of German East Africa is
clothed in imported calico. This was not always the case; even now in
some parts of the north dressed skins are still the prevailing wear,
and in the north-western districts—east and north of Lake
Tanganyika—lies a zone where bark-cloth has not yet been
superseded. Probably not many generations have passed since such
bark fabrics and kilts of skins were the only clothing even in the
south. Even to-day, large quantities of this bright-red or drab
material are still to be found; but if we wish to see it, we must look in
the granaries and on the drying stages inside the native huts, where
it serves less ambitious uses as wrappings for those seeds and fruits
which require to be packed with special care. The salt produced at
Masasi, too, is packed for transport to a distance in large sheets of
bark-cloth. Wherever I found it in any degree possible, I studied the
process of making this cloth. The native requisitioned for the
purpose arrived, carrying a log between two and three yards long and
as thick as his thigh, and nothing else except a curiously-shaped
mallet and the usual long, sharp and pointed knife which all men and
boys wear in a belt at their backs without a sheath—horribile dictu!
[51]
Silently he squats down before me, and with two rapid cuts has
drawn a couple of circles round the log some two yards apart, and
slits the bark lengthwise between them with the point of his knife.
With evident care, he then scrapes off the outer rind all round the
log, so that in a quarter of an hour the inner red layer of the bark
shows up brightly-coloured between the two untouched ends. With
some trouble and much caution, he now loosens the bark at one end,
and opens the cylinder. He then stands up, takes hold of the free
edge with both hands, and turning it inside out, slowly but steadily
pulls it off in one piece. Now comes the troublesome work of
scraping all superfluous particles of outer bark from the outside of
the long, narrow piece of material, while the inner side is carefully
scrutinised for defective spots. At last it is ready for beating. Having
signalled to a friend, who immediately places a bowl of water beside
him, the artificer damps his sheet of bark all over, seizes his mallet,
lays one end of the stuff on the smoothest spot of the log, and
hammers away slowly but continuously. “Very simple!” I think to
myself. “Why, I could do that, too!”—but I am forced to change my
opinions a little later on; for the beating is quite an art, if the fabric is
not to be beaten to pieces. To prevent the breaking of the fibres, the
stuff is several times folded across, so as to interpose several
thicknesses between the mallet and the block. At last the required
state is reached, and the fundi seizes the sheet, still folded, by both
ends, and wrings it out, or calls an assistant to take one end while he
holds the other. The cloth produced in this way is not nearly so fine
and uniform in texture as the famous Uganda bark-cloth, but it is
quite soft, and, above all, cheap.
Now, too, I examine the mallet. My craftsman has been using the
simpler but better form of this implement, a conical block of some
hard wood, its base—the striking surface—being scored across and
across with more or less deeply-cut grooves, and the handle stuck
into a hole in the middle. The other and earlier form of mallet is
shaped in the same way, but the head is fastened by an ingenious
network of bark strips into the split bamboo serving as a handle. The
observation so often made, that ancient customs persist longest in
connection with religious ceremonies and in the life of children, here
finds confirmation. As we shall soon see, bark-cloth is still worn
during the unyago,[52] having been prepared with special solemn
ceremonies; and many a mother, if she has no other garment handy,
will still put her little one into a kilt of bark-cloth, which, after all,
looks better, besides being more in keeping with its African
surroundings, than the ridiculous bit of print from Ulaya.
MAKUA WOMEN

You might also like