Hi Betsy, This is very valuable. Discussed yesterday with Tamara, Dan and Vince. They'll review on the dev call today with some cool work they've done to automate data quality checks by site. I think they will then join your call Wedneday and I will try to make it as well,
Russ From: Chrischilles, Elizabeth A [mailto:[email protected]] Sent: Friday, February 20, 2015 3:45 PM To: Russ Waitman Cc: Theresa Shireman; McDowell, Bradley D; Jianghua He; Tamara McMahon Subject: update on breast cancer sample selection and your assistance needed Dear Russ, Pardon the length of this email. Hope it's worth your read. I've attached graphs of the data elements that we have been able to read for each site as an FYI. As you astutely mentioned on yesterday's GPC Global call, the Breast Cancer Cohort Identification Detectives (AKA Wendy, Theresa, Brad and Tamara) have indeed encountered a substantial roadblock having to do with inconsistencies across the sites. At present only KUMC data allow us to apply the exclusion criteria needed to select the sample. No other site has data that can be read by Wendy. We need to send the ball back to you and the informatics team to get the data ready for use. Let me explain the exploration we've done so far and what the critical elements are: Focusing just on the data needed to select the sample (and ignoring all the other descriptive data elements which also can't be read) we've been able to do a little bit of investigation. In some cases it appears that a data element was never received (a serious problem that we need to go back to the sites on ASAP) and in many cases the data element seems to have been received but not recognized by Wendy's programs. However, the process of gleaning this information (i.e. never submitted vs submitted but in some different form) is excruciating and our group has come to a standstill. Here are the crucial data elements (all from the registries): * Primary site * Sex * Sequence Number * Diagnostic Confirmation * Morphology Code * Derived AJCC-7 Grp and/or SS2000 * Vital Status Ø Only KUMC has complete data for all of the above (at least that Wendy could read) Ø Only KUMC and MC have Sequence Number. Other sites appear to have not sent this data element (at least where we spot checked). One more really important thing - we do know that sequence number was not submitted by most sites and I propose we get the sites working on finding this data element right now. The data request query specified 380 Sequence Number - Central. If the sites did not read that into i2b2 from their tumor registries, they will need to find out if their registries populate that field and, if so, read it in. There is another related field that could be a substitute if the central sequence number is not available: 560 Sequence Number - Hospital. If they have read that variable in we should ask them to provide that now. Lastly, we need to sample at least 200 patients after exclusions so we need at least 285 cases submitted for each site. From the number of records read in we already know that we will need to expand the diagnosis window back to January 1, 2013 for UT San Antonio (only 102 patients submitted), Nebraska (only 240 patients submitted), and Iowa (218 patients submitted). We don't know yet about MCW because it came in late. Let us know what your team can do and how long you think it will take. We'll adjust the survey timeline after we hear that. The current timeline called for the sample to be selected by March 2. Thanks much, Betsy Elizabeth A. Chrischilles, PhD Professor and Marvin A. and Rose Lee Pomerantz Chair in Public Health Department of Epidemiology College of Public Health S424 CPHB 145 N. Riverside Dr. Iowa City, Iowa 52242-2007 Voice: 319-384-1575 Fax: 319-384-4155
_______________________________________________ Gpc-dev mailing list [email protected] http://listserv.kumc.edu/mailman/listinfo/gpc-dev
