Hi Meng-Ying
The calculation of the experimental variance on a finite set of data (population or sample) is simply a mathematical operation
- in itself it has no more meaning that say adding the square of the first value to the cube root of the second and dividing the answer
by the geometic mean of the rest of them.
What bestows meaning on this particular calculation is (roughly speaking)
the assumption that
'each of the 27 values vary about the same mean value with the same distribution of variability at each point'. If we did not believe this - if
for example each of the 27 points sampled completely different phenomena -
then there would be little point in using the variance as a means of
describing the data (spatial data or not..). In other words for the variance to have the sort of meaning that we usually ascribe to it as variation about the mean value - we interpret the observations as realisations of some random
variable and then the calculated variance as estimating the mathematical idealisation of the mathematical variance of the RV.
Likewise, if we interpret the data as being spatial data, we may calculate the experimental variogram and try to interpret it as an estimate of the theoretical variogram of some idealised random function. For a standard variogram to be interpretable we need the first and second moments of the first order increments Z(x+h)-Z(x) to be invariant under translations. Under the more stringent criteria that we may reasonably model our data as
2nd order stationary then the mean of Z(x) is constant everywhere and the variability about the mean is the same at each point - so we can calculate the variance. It is then a theorem, in the context of this model, that the sill is equal to the variance.
outside the context of the stationarity hypothesis, then the variance of the data looses its meaning as variation about a mean value - so is a meaningless
calculation. So, it is hardly a surprise, or concern, that it does not agree
with the sill. The variogram seems to retain its objectivity a bit longer,
until the increments are no longer well modeled as stationary.
For the small populations that you give, well the two calculations (variance and sill) are just numbers. If you try to ascribe meaning to them in the context of stationarity - then their likely variations about some 'true' value
comes into play.
Anyhow, I will be away for the next few days so will miss the end of this
topic (much to the relief of everyone on ai-geostat no doubt!) - but it was
fun - and took me back a good few years (wishing i listened a bit better in matheron's classes on his 'estimating and choosing' book)
Regards
Colin Daly
-----Original Message-----
From: Meng-Ying Li [mailto:[EMAIL PROTECTED]]
Sent: Wed 12/8/2004 9:52 PM
To: Colin Daly
Cc: Digby Millikan; ai-geostats
Subject: RE: [ai-geostats] Re: Sill versus least-squares classical variance estimate
Hi Colin,
What I'm talking about in my example is comparing two descriptive
statistics for this population which consists of 27 data points. No
estimation here is involved, so the thing about confidence interval of
the mean or variance is not of concern here. And it doesn't matter which
model I used in the generator or what parameters I used, since I
re-calculated the population sill and variance after the data are
generated.
Let me state this clear:
(Capitalization indicates highlighting, not speaking tone :p)
1. I generated a POPULATION which is, believe it or not, a series of 27
data.
2. The POPULATION variance, in my example, doesn't match the POPULATION
sill calculated in the POPULATION variogram.
3. So how are we going to estimate the POPULATION variance by the sill in
a SAMPLE, when the sill and the variance in the POPULATION just
doesn't match?
And just a personal opinion, I would like to think geostatistic
theories apply to population of any size, as small as 27, or as large as
1,000,000. If I'm making an example that geostatistics doesn't apply, then
there's something to concern about in this approach.
Meng
On Wed, 8 Dec 2004, Colin Daly wrote:
>
> Hi Meng-Ying
>
> 27 points - you can't really calculate a variogram. With a range of 3 -
> you have about 9 correlation lenghts in the field. So as a crude
> approximation, even the standard deviation on the estimate of the mean
> would be of the order of s.d/sqrt(9) (I vaguely remember trying to get a
> more accurate version of this in the case of a Gaussian RF as an
> exercise in one of Matheron's classes...)
>
> so with s.d = 2.8 (or 2.4 ---similar answers), then standard error is
> 2.8/3=0.9 (approx)
>
> so your confidence interval for the mean would be [m-1.8, m+1.8]
>
> - this is the same order for both the estimate of the sill and for the
> direct estimate of the variance... both are bad
>
> That is for the comparitively easy case of the mean - The situation
> for the variance is even worse - so there is no way that you can
> complain about the quality of the estimate.
>
> I'm not sure if you are suggesting that you should get different
> answers - or that there is some bias involved but to convince yourself
> that there is not repeat your experiment but use a length of 1,000,000
> instead of 27....then at least we would get rid of most of the
> statistical fluctuations - and the estimates should be similar. How are
> you generating the random sequence - is it an AR process or something
> where the variance is known theoretically?
>
> Colin
>
> -----Original Message-----
> From: Meng-Ying Li [mailto:[EMAIL PROTECTED]]
> Sent: Wed 12/8/2004 6:36 PM
> To: Digby Millikan
> Cc: ai-geostats
> Subject: Re: [ai-geostats] Re: Sill versus least-squares classical variance estimate
> Hi Digby and All,
>
> I did a little experiment on the idea that Digby mentioned: The sill will
> estimate the population variance, but found it not true in my experiment:
>
> 1. I generated a set of one-dimentional data with 27 points on regular
> unit spacings, which I'd like to take it as the true, or population
> value. On purpose, I generate the data so it has an influence range of
> three length units.
> 2. I calculated the experimental variogram. Notice that the variogram is
> the population variogram. The sill value is around 2.8.
> 3. But the population variance is 2.39, lower than the sill value.
>
> This confirms my doubt about using sill value as the estimate of
> population variance, since I calculate the variogram and variance based on
> all data points. Please tell me what you think. The data I generated are
> as follows:
>
> 0.056970748
> 0.14520424
> 0.849710204
> 1.650514605
> 1.101666385
> 1.015177986
> 2.150259206
> 2.830780659
> 0.223495817
> -2.47615958
> -3.372697392
> -0.530685611
> 0.786582177
> 0.970673
> 0.674755256
> 0.338461632
> 1.020874834
> 0.410936991
> 1.702892405
> 2.649748012
> 4.290179731
> 3.442015668
> 1.488818953
> 0.862788738
> 0.728709892
> 2.398182914
> 1.522546427
>
>
>
>
>
>
>
> DISCLAIMER:
> This message contains information that may be privileged or confidential and is the property of the Roxar Group. It is intended only for the person to whom it is addressed. If you are not the intended recipient, you are not authorised to read, print, retain, copy, disseminate, distribute, or use this message or any part thereof. If you receive this message in error, please notify the sender immediately and delete all copies of this message.
* By using the ai-geostats mailing list you agree to follow its rules ( see http://www.ai-geostats.org/help_ai-geostats.htm )
* To unsubscribe to ai-geostats, send the following in the subject or in the body (plain text format) of an email message to [EMAIL PROTECTED] Signoff ai-geostats
