Hi Betsy,
This is very valuable.  Discussed yesterday with Tamara, Dan and Vince.   
They'll review on the dev call today with some cool work they've done to 
automate data quality checks by site.  I think they will then join your call 
Wedneday and I will try to make it as well,

Russ


From: Chrischilles, Elizabeth A [mailto:[email protected]]
Sent: Friday, February 20, 2015 3:45 PM
To: Russ Waitman
Cc: Theresa Shireman; McDowell, Bradley D; Jianghua He; Tamara McMahon
Subject: update on breast cancer sample selection and your assistance needed

Dear Russ,

Pardon the length of this email.  Hope it's worth your read.

I've attached graphs of the data elements that we have been able to read for 
each site as an FYI.  As you astutely mentioned on yesterday's GPC Global call, 
the Breast Cancer Cohort Identification Detectives (AKA Wendy, Theresa, Brad 
and Tamara) have indeed encountered a substantial roadblock having to do with 
inconsistencies across the sites.  At present only KUMC data allow us to apply 
the exclusion criteria needed to select the sample.  No other site has data 
that can be read by Wendy.  We need to send the ball back to you and the 
informatics team to get the data ready for use.

Let me explain the exploration we've done so far and what the critical elements 
are:

Focusing just on the data needed to select the sample (and ignoring all the 
other descriptive data elements which also can't be read) we've been able to do 
a little bit of investigation. In some cases it appears that a data element was 
never received (a serious problem that we need to go back to the sites on ASAP) 
and in many cases the data element seems to have been received but not 
recognized by Wendy's programs.  However, the process of gleaning this 
information (i.e. never submitted vs submitted but in some different form) is 
excruciating and our group has come to a standstill.  Here are the crucial data 
elements (all from the registries):
*       Primary site
*       Sex
*       Sequence Number
*       Diagnostic Confirmation
*       Morphology Code
*       Derived AJCC-7 Grp and/or SS2000
*       Vital Status
Ø  Only KUMC has complete data for all of the above (at least that Wendy could 
read)
Ø  Only KUMC and MC have Sequence Number.  Other sites appear to have not sent 
this data element (at least where we spot checked).

One more really important thing - we do know that sequence number was not 
submitted by most sites and I propose we get the sites working on finding this 
data element right now.  The data request query specified 380 Sequence Number - 
Central.  If the sites did not read that into i2b2 from their tumor registries, 
they will need to find out if their registries populate that field and, if so, 
read it in.  There is another related field that could be a substitute if the 
central sequence number is not available: 560 Sequence Number - Hospital. If 
they have read that variable in we should ask them to provide that now.

Lastly, we need to sample at least 200 patients after exclusions so we need at 
least 285 cases submitted for each site.  From the number of records read in we 
already know that we will need to expand the diagnosis window back to January 
1, 2013 for UT San Antonio (only 102 patients submitted), Nebraska (only 240 
patients submitted), and Iowa (218 patients submitted).  We don't know yet 
about MCW because it came in late.

Let us know what your team can do and how long you think it will take.  We'll 
adjust the survey timeline after we hear that.  The current timeline called for 
the sample to be selected by March 2.

Thanks much,
Betsy

Elizabeth A. Chrischilles, PhD
Professor and Marvin A. and Rose Lee Pomerantz Chair in Public Health
Department of Epidemiology
College of Public Health
S424 CPHB
145 N. Riverside Dr.
Iowa City, Iowa 52242-2007

Voice: 319-384-1575
Fax: 319-384-4155

_______________________________________________
Gpc-dev mailing list
[email protected]
http://listserv.kumc.edu/mailman/listinfo/gpc-dev

Reply via email to