Hi.

On Thu, Jun 18, 2009 at 9:44 AM, han<[email protected]> wrote:
>
> Hi Henrik:
>
> I am looking at SNP 500K data for cell-lines that was released from
> GSK.

Are you referring to the 'GSK Cancer Cell Line Genomic Profiling Data'
data set listed on:

http://groups.google.com/group/aroma-affymetrix/web/data-sets?

I will assume that.

> I followed your example CRMA v2.  with CBS segmentation without
> reference samples.  It seems that the scale/dynamic range in the plus
> side (amplification) is compressed compared with similar data from
> other sources (same CEL file, or same cell-lines).  For example, the
> erbbB/Her2 region (chr 17, ~35 mb) in BT474 cell-line, I got relative
> copy number ~ 2, which translate to 4 copies.  The same region for
> BT474 from Sanger center (their CEL data) has ~ 15 copies, while NCI
> cancer genome workbench (same CEL file, not sure what algoruthm) shows
> ~28 copies.
>
> I realized that many factors could impact copy number calculation,
> just want to know your thoughts on this.  Could it duo to using
> average as reference? or something specific for your algorithm?

So, since this is cell-line data one should be able to safely assume
that the "purity" of these are the same regardless how ran the
experiments.  Otherwise, things like normal contamination will
naturally shrink the signals toward the normal, regardless whether you
think in terms of total or allele-specific CNs.

Out ruling this factor here, the most common reason for observing
different compression factors is how much offset is subtracted from
the raw signals - think M = log2(C), C = (tumor + a) / (normal + a)
and adjust a.  Thus, if you don't subtract enough offset, your CN
ratios will be compress.  The compression actually depends on
affinities etc so it will vary from locus to locus, cf. [1].  It is
clear that different preprocessing methods adjust for different amount
of offset, e.g. dChip and CRMA v1/2 tend to correct for more than, say
CNAT/CN5.  If you look at Figures 6 & 7 in the CRMAv2 paper [2], you
can (somewhat) see this from the CN estimates.  However, from you data
this would mean that CRMA v2 subtracts less offset than the other
(unknown) methods.  Without knowing the method, I'm not sure this is
the case because traditionally method developers have been reluctant
to do this in order to avoid getting too close to theta=0 or below (we
accept that!).

Another alternative is that NCI workbench is correcting for the tumor
(in-)purity, that is, that they remove the normal component from the
signal before calculating the ratios.  That would bring up the
*observed* CN levels.

Please note that we do not claim to *calibrate* the observed CN levels
in CRMA v1/2, e.g. by subtract normal contamination etc.  This is
deliberate, because we think that is a downstream task.  I cannot
speak for other methods, but I'm sure that's a common strategy.
However, we claim that the CN rank should be (approximately) the same.
 In other words, to pay to much attention to the absolute level of the
CN aberrations, but only to their relative levels (and where there
breakpoints are).  Related to this, in our study for combining CN
estimates from different labs and different technologies [3], we found
that although the different sources provide CN ratios that are
differently compressed, the all show the same profiles.  After
normalizing for the scale differences, they all agree a lot, cf.
Figure 2 and Figure 5.

To summarize, I conclude that this is mostly a calibration issue, and
that you need to be aware what the preprocessing method do and how
they work to interpret and compare absolute CN levels across
methods/labs.

Hope this helps

Henrik

REFERENCES (all open access):
[1] H. Bengtsson and O. Hössjer, Methodological study of affine
transformations of gene expression data with proposed robust
non-parametric multi-dimensional normalization method, BMC
Bioinformatics, 2006, 7:100.
http://www.biomedcentral.com/1471-2105/7/100/

[2] H. Bengtsson; P. Wirapati & T.P. Speed, A single-array
preprocessing method for estimating full-resolution raw copy numbers
from all Affymetrix genotyping arrays including GenomeWideSNP 5 & 6,
Bioinformatics, 2009.
http://bioinformatics.oxfordjournals.org/cgi/content/short/btp371v1

[3] H. Bengtsson, A. Ray, P. Spellman & T.P. Speed, A single-sample
method for normalizing and combining full-resolution copy numbers from
multiple platforms, labs and analysis methods, Bioinformatics, 2009.
[pmid: 19193730] [doi: 10.1093/bioinformatics/btp074]
http://bioinformatics.oxfordjournals.org/cgi/content/short/25/7/861

>
> Thanks a lot.
>
> Han Chang
> >
>

--~--~---------~--~----~------------~-------~--~----~
When reporting problems on aroma.affymetrix, make sure 1) to run the latest 
version of the package, 2) to report the output of sessionInfo() and 
traceback(), and 3) to post a complete code example.


You received this message because you are subscribed to the Google Groups 
"aroma.affymetrix" group.
To post to this group, send email to [email protected]
To unsubscribe from this group, send email to 
[email protected]
For more options, visit this group at 
http://groups.google.com/group/aroma-affymetrix?hl=en
-~----------~----~----~----~------~----~------~--~---

Reply via email to