Among number of challenges faced by informatics one of the long standing and critical challenge has been the Biodiversity informatics: the challenge of linking data and the role of shared identifiers.
A major challenge facing biodiversity informatics is integrating data stored in widely distributed databases. Initial efforts have relied on taxonomic names as the shared identifier linking records in different databases. However, taxonomic names have limitations as identifiers, being neither stable nor globally unique, and the pace of molecular taxonomic and phylogenetic research means that a lot of information in public sequence databases is not linked to formal taxonomic names. This review explores the use of other identifiers, such as specimen codes and GenBank accession numbers, to link otherwise disconnected facts in different databases. The structure of these links can also be exploited using the PageRank algorithm to rank the results of searches on biodiversity databases. The key to rich integration is a commitment to deploy and reuse globally unique, shared identifiers [such as Digital Object Identifiers (DOIs) and Life Science Identifiers (LSIDs)], and the implementation of services that link those identifiers.
Showing posts with label controlled vocabularies. Show all posts
Showing posts with label controlled vocabularies. Show all posts
Tuesday, April 29, 2008
Yet another challenge to informatics, well does it have the answer this time?
Labels:
bio software,
bioinformatics,
bioit,
controlled vocabularies,
curate,
database,
digitized,
DOI,
extracted,
identifiers,
informatics,
knowledge integration,
LSID,
ontology,
Semantic Web,
taxonomy
Thursday, April 24, 2008
Structured Digital Abstracts - Easier Literature Searching
Already blogging on similar lines and the subject of Bring in data from the published literature in my earlier blog Smart tools to track, analyze and visualize research , here is another interesting experiment on similar lines:
The experiment centres on Structured Digital Abstracts (SDA). SDA are extensions of the normal journal article abstracts that describe the relationship between two biological entities, mentioning the method used to study the relationship. Each sentence is preceded by one or more identifiers pointing to the corresponding database entries that contain the full details of the interaction e.g. protein A interacts with protein B, by method X.
The aim of SDA is to assist data entry, text mining and literature searching by extracting the salient data from the article into simple sentences using a defined structure and controlled vocabularies.
The experiment centres on Structured Digital Abstracts (SDA). SDA are extensions of the normal journal article abstracts that describe the relationship between two biological entities, mentioning the method used to study the relationship. Each sentence is preceded by one or more identifiers pointing to the corresponding database entries that contain the full details of the interaction e.g. protein A interacts with protein B, by method X.
The aim of SDA is to assist data entry, text mining and literature searching by extracting the salient data from the article into simple sentences using a defined structure and controlled vocabularies.
Gianni Cesareni, Editor of FEBS Letters explains:
Many articles in biological journals describe relationships between entities (genes, proteins, etc.) yet this information cannot be efficiently used because of difficulties in retrieving from text. Databases capture this valuable information and organize it in a structured format ready for automatic analysis. The experiment of using SDAs will facilitate database entry and improve disclosure, to the benefit of authors and readers.
Labels:
bio software,
controlled vocabularies,
curation,
database,
digitized,
entities,
extracted,
informatics,
literature mining,
NLP,
protein interaction,
SDA,
Structured Digital Abstracts,
text mining
Subscribe to:
Posts (Atom)