As stated in part II of this post sequence, we initiate the development of BRIL based on the assumption that five elements can capture the majority of resources used by biomedical researchers to evaluate whether a collaboration is warranted: databases, analysis methods, infrastructure (e.g., clinics, software, etc), staff (e.g., research coordinators, lab technicians, etc), and experts (e.g., publications, education, funding). In this post we will describe the elements related to databases and data analysis methods and how some of the applications developed in the Center for Excellence in Surgical Outcomes attempted to capture the data associated with these elements. Of importance, each of the described applications does not contain an underlying BRIL ontology, but the BRIL ontology has the potential to link all of these applications, a topic that will be described in a future post focusing on use cases. So, to the elements and respective applications:
1. Database
The idea behind the concept of providing information about the databases resource was initially launched by our group using the concept of a databases of databases (Pietrobon, 2004). The idea is that the database resource can be described using two levels.
a. First, databases should be first generally described using elements such as title, keywords, contact information for database owner or director, overall description, total number of observations, age group of population, data collection method, type of data, geographical area, clinical area, variables that are carried longitudinally, clinical systems, time period, statistical sampling method, suggestions for research questions, and previous publications making use of database.
b. Second, a data dictionary is provided containing the original variable name in the database, its corresponding question, alternative responses, and general measures of variable quality (e.g., missing rate). The initial version of the database of databases allowed for searches across these fields, while the current second version also allows for modifications of individual fields by regular users through a wiki-like mechanism.
At this time, the database of databases has been upgraded to a system that not only allows for searching of information on several clinical databases, but also allow for immediate modification of its information (Wikimedic application).
2. Data analysis methods
The idea behind the concept of providing information about data analysis methods was also launched by our group using the concept of a layers of information (Pietrobon, 2004). The idea is that the description of data analysis methods can be simplified using five levels or layers of information.
a. Layer 1 or general description: This layer is the simplest of all, simply providing a one-paragraph, lay description of what the data method analysis is and what can be accomplished with it. What we have observed over time,is that, although this layer was initially described as a separate element, it tends to be merged with the layer on annotated references (layer 5). In other words, the Web now has such a wealth of examples describing most data analysis methods, that it is usually easy to find a simple, down to earth explanation about a method that will satisfy most researchers being introduced to that method.
b. Layer 2 or previous examples of the data analysis method used in previous biomedical research publications: The basis for this layer is that most adult learning occurs based on the establishment of a connection with previous knowledge. In the case of a clinical researchers, for example, it is simpler for them to read a published article in their fields using a new data analysis method than to read an article describing a series of theoretical assumptions about the statistical underpinnings of the same method. In other words, if they start from something familiar (a clinical scenario) and only then go to something new (a data analysis method used to clarify a research question in that clinical scenario), the learning process is made easier. In practice, the layer for previous examples is simply a group of links to peer-reviewed articles using the analysis method under scrutiny, usually giving preference to open access articles so that they can be freely accessible by researchers in different parts of the globe.
c. Layer 3 or data input specification: This is probably one of the most technical layers, in that it specifies what types of variables are required to use the analysis method. For example, in order to use the analysis method called "1PL item response theory model," the necessary data input are variables that represent a latent construct, and whose alternative responses are dichotomous (yes/no). This is a more technical layer, but combined with the other layers, it assists researchers and statistical programmers in determining whether an existing database has the necessary variables to use a data analysis method. In addition, if planning a prospective study, this layers allows researchers to determine what type of variables should be collected in order to perform a certain analysis after the study is completed. Finally, this layer is essential in establishing the connection between, for example, a research group that owns a certain database and another group that has expertise with a certain data analysis methodology. This examples will later be expanded in the post on use cases for BRIL.
d. Layer 4 or examples of data output: This layer presents tables and graphics from existing articles using the method. The idea is that researchers can have a good grasp on what to expect as the output to be obtained after the data analysis method is applied. Although this might seem simplistic, researchers trying to evaluate a certain analysis method many times are left wondering exactly what they would get after using a certain methodology. For example, simply going over a technical paper describing the basics of survival analysis does not necessarily give researchers the idea that their final output could be a Kaplan Meier plot, a table with mean survival times, or a table with regression coefficients from a Cox Proportional Hazards model. Similar to the layer with previously published articles using the analysis method, this layer aims at demonstrating to researchers exactly what they will get if they decide to use the method.
e. Layer 5 or annotated references: As mentioned before, the layers containing the general description of the method and the layer with annotated references tend to be consolidated into a single layer. Simply stated, the aim of this layer is to provider researchers and statistical programmers with a large body of literature on the topic, going from the most simple to the most complex. From a biomedical researcher point of view, equations usually mean complexity, and so the simplest references tend to be math-free, usually drawing from rhetoric mechanisms such as analogies and metaphors that simplify reasoning, as well as using plenty of graphics and flowcharts to simplify understanding. More complex explanations will usually draw heavily from mathematics, being specific to certain areas of the method. All references in this layers will usually have a line or two describing their main content, complexity level, and target audience.
How are layers of information implemented? Since they are simply a group of texts, Wiki sites are usually the best way to make them available and constantly updated by researchers and statistical programmers.
Subscribe to:
Post Comments (Atom)
No comments:
Post a Comment