Hi. On Thu, Jun 18, 2009 at 9:44 AM, han<[email protected]> wrote: > > Hi Henrik: > > I am looking at SNP 500K data for cell-lines that was released from > GSK.
Are you referring to the 'GSK Cancer Cell Line Genomic Profiling Data' data set listed on: http://groups.google.com/group/aroma-affymetrix/web/data-sets? I will assume that. > I followed your example CRMA v2. with CBS segmentation without > reference samples. It seems that the scale/dynamic range in the plus > side (amplification) is compressed compared with similar data from > other sources (same CEL file, or same cell-lines). For example, the > erbbB/Her2 region (chr 17, ~35 mb) in BT474 cell-line, I got relative > copy number ~ 2, which translate to 4 copies. The same region for > BT474 from Sanger center (their CEL data) has ~ 15 copies, while NCI > cancer genome workbench (same CEL file, not sure what algoruthm) shows > ~28 copies. > > I realized that many factors could impact copy number calculation, > just want to know your thoughts on this. Could it duo to using > average as reference? or something specific for your algorithm? So, since this is cell-line data one should be able to safely assume that the "purity" of these are the same regardless how ran the experiments. Otherwise, things like normal contamination will naturally shrink the signals toward the normal, regardless whether you think in terms of total or allele-specific CNs. Out ruling this factor here, the most common reason for observing different compression factors is how much offset is subtracted from the raw signals - think M = log2(C), C = (tumor + a) / (normal + a) and adjust a. Thus, if you don't subtract enough offset, your CN ratios will be compress. The compression actually depends on affinities etc so it will vary from locus to locus, cf. [1]. It is clear that different preprocessing methods adjust for different amount of offset, e.g. dChip and CRMA v1/2 tend to correct for more than, say CNAT/CN5. If you look at Figures 6 & 7 in the CRMAv2 paper [2], you can (somewhat) see this from the CN estimates. However, from you data this would mean that CRMA v2 subtracts less offset than the other (unknown) methods. Without knowing the method, I'm not sure this is the case because traditionally method developers have been reluctant to do this in order to avoid getting too close to theta=0 or below (we accept that!). Another alternative is that NCI workbench is correcting for the tumor (in-)purity, that is, that they remove the normal component from the signal before calculating the ratios. That would bring up the *observed* CN levels. Please note that we do not claim to *calibrate* the observed CN levels in CRMA v1/2, e.g. by subtract normal contamination etc. This is deliberate, because we think that is a downstream task. I cannot speak for other methods, but I'm sure that's a common strategy. However, we claim that the CN rank should be (approximately) the same. In other words, to pay to much attention to the absolute level of the CN aberrations, but only to their relative levels (and where there breakpoints are). Related to this, in our study for combining CN estimates from different labs and different technologies [3], we found that although the different sources provide CN ratios that are differently compressed, the all show the same profiles. After normalizing for the scale differences, they all agree a lot, cf. Figure 2 and Figure 5. To summarize, I conclude that this is mostly a calibration issue, and that you need to be aware what the preprocessing method do and how they work to interpret and compare absolute CN levels across methods/labs. Hope this helps Henrik REFERENCES (all open access): [1] H. Bengtsson and O. Hössjer, Methodological study of affine transformations of gene expression data with proposed robust non-parametric multi-dimensional normalization method, BMC Bioinformatics, 2006, 7:100. http://www.biomedcentral.com/1471-2105/7/100/ [2] H. Bengtsson; P. Wirapati & T.P. Speed, A single-array preprocessing method for estimating full-resolution raw copy numbers from all Affymetrix genotyping arrays including GenomeWideSNP 5 & 6, Bioinformatics, 2009. http://bioinformatics.oxfordjournals.org/cgi/content/short/btp371v1 [3] H. Bengtsson, A. Ray, P. Spellman & T.P. Speed, A single-sample method for normalizing and combining full-resolution copy numbers from multiple platforms, labs and analysis methods, Bioinformatics, 2009. [pmid: 19193730] [doi: 10.1093/bioinformatics/btp074] http://bioinformatics.oxfordjournals.org/cgi/content/short/25/7/861 > > Thanks a lot. > > Han Chang > > > --~--~---------~--~----~------------~-------~--~----~ When reporting problems on aroma.affymetrix, make sure 1) to run the latest version of the package, 2) to report the output of sessionInfo() and traceback(), and 3) to post a complete code example. You received this message because you are subscribed to the Google Groups "aroma.affymetrix" group. To post to this group, send email to [email protected] To unsubscribe from this group, send email to [email protected] For more options, visit this group at http://groups.google.com/group/aroma-affymetrix?hl=en -~----------~----~----~----~------~----~------~--~---
