[With some reformatting (line, paragraph breaks) that, hopefully,
altered neither content nor intent. -the Moderator]
Dear Morphmet colleagues,
Dennis Slice has invited me to comment on the article by Douglas
Theobald and Deborah Wuttke published as PNAS 103:18521-18527 and posted
to this news group (via a link) on February 5. I am happy to do so.
These remarks are not intended to foreclose comments by others,
responses by Theobald and Wuttke, or any other form of continuing
discussion. In the text below, purely to save space, I will use the
acronym "TW" to mean either "Drs. Theobald and Wuttke" or "the paper by
Theobald and Wuttke," whichever is appropriate for the sentence
structure.
I apologize in advance for the length of these remarks. While
some arise from a close reading of TW, most are redacted from a
fifty-page manuscript chapter that will circumscribe the core of the
long-delayed revision of the Orange Book if there ever is to be one. Now
that we have a couple of decades' experience in successful applications
of geometric morphometrics (GMM), it is time to revisit foundations.
The TW paper begins (in a statement I will disagree with below):
The goal of Procrustes analysis is to superpose
(*) non-identical shapes in an optimal manner via scalings,
translations, and rotations.
They note that the ordinary least-squares [LS] superposition is
not an optimal estimator of anything under most more general assumptions
about the distribution of actual landmark locations in the real
Euclidean space of the raw coordinate data. (When the landmark data are
distributed according to the so-called "offset isotropic normal model,"
the LS estimate is in fact ML.) They set out an alternative, a
maximum-likelihood [ML] solution, under one such more general assumption
whereby landmark-by-landmark deviations from the mean are anisotropic
but also correlated in one particular patterned way.
My response weaves three threads together. They were difficult
to separate, though, and I would suggest that the reader chew all the
way to the end before beginning to think critically about anything in
the middle.
1. The concern for an optimal superposition arises mainly in
sciences for which such a superposition has a scientific meaning that is
the subject of explanation or verification at other levels. Applications
of GMM in molecular and structural biology may come under this heading,
although I think they would need to be in the version of the analysis
that involves form space, not shape space. But applications in
developmental biology, evolutionary biology, and medicine seem
uninformed by those questions. In these domains,
2. The least-squares property of the usual Procrustes
superpositions has nothing to do with any aspect of real data, but
arises from the geometry of the Kendall quotient space within which our
scientific descriptions of shape or form now reside. The least-squares
formulas express some symmetries of our preferred descriptive style;
they are not associated with any particular protocol of estimation of
anything.
3. The sentence (*) from TW does not actually match its use in
the scientific context sketched in (2). In those sciences, Procrustes
analysis is the multivariate analysis of one of two different Procrustes
distances -- either Procrustes shape distance or Procrustes form
distance -- and the role of the superpositions is to supply a set of
variables that account for the appropriate sums of squares. It is the
least-squares superpositions, and only the least-squares superpositions,
that do so; covariances already implicit in the distance matrix are
reconstituted from these superpositions, not modeled in advance. Hence
the superposition is fixed by the rules of the game; it is not the topic
of any estimation protocol.
Let me expand on each of these points in turn.
1. TW's Gaussian perturbation model is a special case of the
general Mardia-Dryden model according to which the general covariance
structure of that model takes a factored form. For assumptions about
this covariance structure to make scientific sense, it must be the case
that some scientific theory is available that makes a claim of this
sort, a restricted model such as this, that can then be checked in
particular data sets.
The question naturally arises as to the nature of a theory that
would account for a covariance matrix of factored form, \Omega = \Sigma
\otimes \Xi in the TW Kronecker notation as if typeset in TeX. Even for
TW's preferred application to "atomic positions," for ML estimation to
be of scientific use, the model the likelihood of which is being
maximized needs to be, at least approximately, TRUE. Hence if "real"
(physical) Procrustes analysis is to have scientific applications, there
must be a nesting of theories somewhere, under one of which the
covariance matrix is factored, against a general alternative permitting
falsification. I have seen no such scientific work, but structural
biology is not my field; there may be examples there.
_Additional note._ In my opinion those molecular studies would
actually require Procrustes form space, not Procrustes shape space.
Whenever one begins from real perturbations of locations in real
Euclidean three-space size information is likely to prove a crucial part
of the analysis, and must be preserved. The radii of molecular orbitals
don't scale; it may follow that the scaled Procrustes superpositions of
TW are not going to be theoretically informed no matter what version of
superposition is used. Our Procrustes form coordinates, see (3) below,
are not the Cartesian coordinates of an unscaled superposition; they are
generated differently. [Squared Procrustes form distance equals squared
Procrustes shape distance plus the square of the log of the ratio of
Centroid Sizes.] TW state in their posting that their intended
scientific context is structural biology, and maybe somebody should
cross-post this comment to those bulletin boards.
2. Procrustes superpositions are least-squares because the
formulation of shape distance in the original Kendall context is
least-squares in the corresponding Riemannian geometry. The TW comment
has conflated two separate uses of the notion of least squares, one
having to do with geometry, one with estimation. Kendall's
least-squares is the _geometric_ usage, the closest (= least-squares)
approach between two orbits in an elliptic Riemannian manifold. This
geometry has nothing to do with multivariate statistics, and the only
probability theory it captures is the theory of isotropic diffusion that
led, via stochastic calculus, to Kendall's original formulations of the
uniform shape distribution (itself a very powerful form of null
hypothesis, but corresponding to entirely different sciences, like
astronomy or archaeology, in which there is no equivalent of a "mean"
form at all). Procrustes _distance_ is defined as the (normalized)
minimum Euclidean distance between orbits over Euclidean space; the
least-squares property is built into the explicit approximation by
sums-of-squares in the quotient space we all use.
Thus the statement (*) from TW with which I opened is not
applicable to the scientific contexts in which I and (I assume) most of
my present readers are working. Procrustes analysis really has nothing
much to do with "finding the optimal superposition," if by optimal is
meant some reference to a real-world phenomenon. Our superpositions are
instead formalisms beginning with the proposition that the
scientifically meaningful relationships between forms are characterized
only up to the operations of the isometry group (form space) or the
similarity group (shape space). The superposition is a computational
device to get at a set of variables (the shape coordinates, with or
without log Centroid Size) dual to these distances (see below). If the
superposition is itself of scientific interest, almost certainly
Procrustes distance is already the wrong scientific descriptor, and one
must revert to far more detailed models and far more detailed empirical
knowledge.
[Note: the log transform of "log Centroid Size" a few lines
above is not really the same as the log of the lognormal distribution in
TW's equation 10. In GMM the log is there for geometrical reasons, not
stochastic reasons. For the reasoning, see Mitteroecker et al., J. Human
Evolution 46:679--697, 2004, the same place where Procrustes form space
was first introduced. This logarithm, like the Procrustes use of
"least-squares," is an expression of the semantics of a description, not
any sort of empirical claim, hence not the object of any "estimation."]
3. If the TW statement (*) about "the goal of Procrustes
analysis" is in fact inappropriate in the non-structural sciences, as I
have been arguing, what gets to replace it? Here's the short version:
Procrustes analysis is the explicit geometric construction of
shape or form coordinates whose sums of squares approximate the
actual Procrustes distances observed among a sample of forms.
If the Procrustes distance used is the conventional "shape
distance," invariant vs. the similarity group, these are
conveniently taken as the locations of the corresponding
Procrustes shape coordinates in superposition on the sample
average shape. If the Procrustes distance is our new "form
distance," invariant vs. the isometry group, then there is one
additional coordinate, the log of Centroid Size.
In other words, the purpose of Procrustes superposition is to
supply the coordinates that reproduce the principal COORDINATES analysis
of a sample of forms when one or the other Procrustes metric is used as
the corresponding distance function. The first principal COMPONENT of a
set of Procrustes shape coordinates is the first principal COORDINATE of
their distances, the second is the second, and so on. We use these
coordinates precisely so that we can duplicate the principal coordinates
as principal components in this way; THAT IS THEIR MAIN JUSTIFICATION.
As a methodology this is a whole lot different from methods in
the sciences that TW are referring to. Here are at least three contrasts.
(i) As John Kent commented a few years ago, in our
applications of GMM the mean form itself is a nuisance parameter right
up there with the similarity transformations. (The comment arose in
context of an argument about "bias" of the Procrustes mean form; Kent
was noting that any such "bias" was utterly without scientific relevance
-- had no implications for any of the real descriptions at which we were
aiming.) In all of the paleoanthropological applications, for instance,
explanations of the mean form are utterly obscured in half a billion
years of vertebrate evolutionary history. It is only the structure of
covariances that expresses processes of growth or selection that are
within our grasp to understand. In the structural biological sciences,
by contrast, sometimes the mean form is capable of being (exactly,
quantitatively) explained. This may be the deepest difference between
the two domains of application.
(ii) In GMM, landmark locations are passive: they are just
moved around by deformations of the covering place or space that is the
domain of actual processes. This is not at all the case in structural
biology, where the applicable physical fields are actually centered at
the landmark locations. The atomic locations, then, parameterize the
true process equations; this is unlikely to be the case in organismal
applications. The nature of quantum phenomena in the spaces in-between,
while probabilistic, is not remotely the kind of probability that
applies in the Procrustes world.
(iii) Molecular objects have real, physical energy that
constrains their form. GMM has no role for any such energy -- the
"bending energy" that the TPS is minimizing is completely metaphorical,
and Procrustes distance itself is a perceptual or semantic metaphor not
claimed to convey any particular biological reality in the applications
known to me. [By way of continuation of a discussion with Paul
O'Higgins:] GMM by itself may be unsuited to biomechanical studies of
physical strain. If finite-element diagrams are the visualization, it
must be their energy term that is the distance metric, not Procrustes
distance. [And, also, because changes of scale entail a cost in elastic
energy, the analysis must be in form space.] The relation between
finite-element scaling analysis [FESA] and Procrustes analysis is
reducible to a relative eigenanalysis of energy terms just like the
relation between bending energy and Procrustes analysis was (the
relative eigenanalysis that gives us the principal and partial warps).
In this multivariate data analysis, which interprets
Procrustes shape coordinates as a graphic for the principal coordinates,
the covariances of the landmark coordinates are never modeled. It is
only the principal coordinate axes themselves that are modeled, and the
modeling is geometric, not statistical: the spatial structure (as
conveyed by thin-plate spline) of the covarying shifts of these shape
coordinates (along with size, for the form-space version). This general
covariance structure is of little interest unless accompanied by the
complementary ordination of specimens in their shape space or form
space: GMM is at root one big application of biplots with some elegant
additional diagrams (the splines). Specific aspects of the covariance
structure are reported in particular settings: symmetry, for instance
(which is, in a sense, the "opposite" of a factored covariance
structure), allometry (strong correlations of all the shape coordinates
with the last column of the form-coordinates matrix), decompositions
into principal warps (large-scale factors) or intrinsic relative warps
(scale-free factors), and others. Hence TW's comment (page 18525) that
their improved superposition coordinates should improve estimation of
principal components has our applications precisely backward. We
already have our principal component scores -- they are principal
coordinates of Procrustes distance. What we need are the variables for
which they are the principal components, and that is what the
least-squares superposition gives us.
That the TW model leads to a feasible algorithm, then, offers
us no particular scientific advantages without a great deal of auxiliary
argumentation and testing alongside quantitative theory far stronger
than anything we can realistically expect to encounter in organismal
biology. The models we most often adduce in real data cannot have a
factored covariance structure -- these certainly do not apply to
semilandmarks, for instance, which is where most cutting-edge GMM work
is centered these days -- and the TW approach (or, for that matter,
appropriately constrained estimation under the general alternative
covariance structure of Mardia and Dryden) would lead to almost
identical data analyses at the final stage, principal coordinates
analysis followed by thin-plate spline (the current central GMM data
flow). The TW method would prove superior in those sciences where the
concern was never with the Procrustes distances per se but instead with
models of landmark positioning in the covering space, e.g. molecular
models or finite-element models for which the mean form is itself of
empirical interest. But, still, sound data-analysis practice in those
sciences would need to nest the factored covariance model under a more
general alternative, precisely so that the data would afford the
possibility of testing that assumption. In today's GMM, exactly the
same scientific purpose, though not in the same notation, is achieved
via the principal components of the shape coordinates (which is to say,
the principal coordinates of the original data matrix of Procrustes
shape distances or Procrustes form distances).
The least-squares tool that Theobald and Wuttke propose to
generalize arises at root as a rule of method for descriptors, not as a
statement about data but as a semantics in which such statements will be
expressed. Their method will prove an advance precisely to the extent
that we can find scientific contexts in which the factored covariance
model in the covering space is true, so that the additional descriptors
\Sigma and \Xi are worth bothering with. In organismal biology, they
seem not to be. Not even the superpositions have any meaning, only the
distances and ordinations that follow algebraically.
In closing, I will return to TW's original words, and extract
two other statements from their introduction that seem to me to be worth
commenting on.
(i) TW state, "[D]ifferent landmarks can have widely different
variances. Under these conditions, ordinary LS can give misleading
results." The statement is itself misleading inasmuch as the Procrustes
superpositions at which TW drive are not the "results" of a Procrustes
data analysis. Our actual results are the principal coordinates of the
sample ( = the principal components of the shape coordinates) along with
the thin-plate splines visualizing them as deformations. These
illustrations are not "misleading" to the extent that the topic is in
fact the ordination of Procrustes distances. They do not and never were
intended to inform about superposition-dependent aspects of covariances
in the covering space.
(ii) TW state, "The method of maximum likelihood is a common
alternative to LS and is widely considered to be fundamental for
statistical modeling and parameter estimation." True. But contemporary
GMM is not a methodology for either purpose: neither for statistical
modeling nor for parameter estimation. It is exploratory multivariate
data analysis aimed at graphical and geometric interpretations of
samples of forms for the purpose of generating hypotheses for subsequent
exploration by other, preferably more directed methods (knockout
experiments, ecophenotypic explorations, evo-devo studies). If we were
doing statistical modeling, we would have calibration data for the
landmark-specific covariances that GMM shoves over onto the principal
coordinates. If we were doing parameter estimations, we would know what
the list of real parameters was -- in organismal biology, it is
certainly not the list of landmark mean locations, which are an accident
of the perceptual psychology by which we characterize certain loci as
"landmarks" -- and in that case we would have no particular use for the
Procrustes distances except as some sort of model of the scientist's own
perceptual system.
Instead, in the overwhelming majority of GMM applications known
to me, there exist neither real statistical models nor real parameters
of anything except principal coordinates of Procrustes distances. In
this context, GMM is a robust and reliable tool for visualizing patterns
in empirical data whenever the underlying formalism of Procrustes
distance (quotiented Euclidean distance) is not in violation of any
prior knowledge. TW's application require there to be such prior
knowledge -- which is possible -- but also that its modeling via a
factored covariance in the covering space be true. That is likely to
prove an insuperable barrier to truth in most otherwise useful
applications, but parts of structural biology just might be domains of
exception to this dogma.
Fred Bookstein
Seattle, Washington, and Vienna, Austria
February 8, 2007
[EMAIL PROTECTED]
--
Replies will be sent to the list.
For more information visit http://www.morphometrics.org