Showing posts with label curate. Show all posts
Showing posts with label curate. Show all posts

Sunday, March 18, 2012

Microbial Genomes Curator @ Computercraft Corporation--Maryland (US)

Microbial Genomes Curator @ Computercraft Corporation--Maryland (US). Submitted by Computercraft Corporation; posted on Saturday, March 17, 2012

RESPONSIBILITIES:
Computercraft seeks a microbiologist to work with a team of software developers and biologists on microbial genome analysis including pan-genome, protein clusters, phylogenetic tree and more. This is a technically challenging position requiring experience in genome sequencing and annotation. A background in comparative genome analysis such as alignments and tree building is a plus.

Our scientists work with genomic experts at the NIH's National Center for Biotechnology Information (NCBI) to create and enhance a suite of databases and tools available to researchers worldwide. Teamwork interaction and excellent organizational skills are essential for this detail-oriented position, as is scientific problem-solving with a results-oriented focus.

REQUIREMENTS:
* PhD in molecular biology, microbiology, or related field
* Experience in genome sequencing and annotation
* Familiarity with BLAST, genome browser, and genome assembly data
* Excellent verbal and written communication skills as well as organizational skills
* Strong interest in contributing to the development of public database resources

PREFERENCES:
Other Desirable Skills:
* Programming experience with LINUX/UNIX
* Scripting experience in PERL or related scripting languages

COMPENSATION:
Computercraft offers a competitive salary and an excellent benefits package including PPO health insurance with 100% company paid premiums, 401K program with matching, paid time off and holiday pay, life insurance, flexible spending and disability coverage. We offer an excellent work life balance with a standard 40 hour work week and the chance to work alongside accomplished scientists at NIH/NCBI.

HOW TO APPLY:
To apply for this position or learn about other Computercraft job opportunities, please visit the Careers section of our website: http://www.computercraft-usa.com/

POLICY:
Computercraft is an equal opportunity employer.

Tuesday, August 11, 2009

Curated databases and data curation

"There does appear to be a distinction between the way curation is used in the bio-sciences, and elsewhere. In particular, the term "curated database" tends to mean a manually constructed database that links literature to data, curated by experts who provide authority (eg see the Wikipedia definition of Biocurator). The earliest mention of the term "curated database" I can find is in the abstract (and only in the abstract) of Larsen et al (1993)." Chris Rusbridge Digital Curation Blog
Since these database are hand curated by experts (manually curated), they always promise a accuracy & quality better than uncurated or NLP based databases. While NLP based databases follow a automated curation provide quick updates and tend to be large in terms of the quantum of data. While they may trade off in accuracy due to their automated curation process.

While platforms like XTractor Premium follow a unique approach by trying to adopt the best of both worlds. A hybrid approach of a first level of NLP which promises a quick turnaround and large quantum of data followed by expert curation to preserve and maintain the high quality promised by expert curated database. XTractor Premium is a specialized biomedical text mining platform with semantically enriched search (Semantic search) and analytics to that enable discovery, knowledge sharing, analysis and modelling of published biomedical facts.
















Tuesday, April 29, 2008

Yet another challenge to informatics, well does it have the answer this time?

Among number of challenges faced by informatics one of the long standing and critical challenge has been the Biodiversity informatics: the challenge of linking data and the role of shared identifiers.

A major challenge facing biodiversity informatics is integrating data stored in widely distributed databases. Initial efforts have relied on taxonomic names as the shared identifier linking records in different databases. However, taxonomic names have limitations as identifiers, being neither stable nor globally unique, and the pace of molecular taxonomic and phylogenetic research means that a lot of information in public sequence databases is not linked to formal taxonomic names. This review explores the use of other identifiers, such as specimen codes and GenBank accession numbers, to link otherwise disconnected facts in different databases. The structure of these links can also be exploited using the PageRank algorithm to rank the results of searches on biodiversity databases. The key to rich integration is a commitment to deploy and reuse globally unique, shared identifiers [such as Digital Object Identifiers (DOIs) and Life Science Identifiers (LSIDs)], and the implementation of services that link those identifiers.

Tuesday, April 22, 2008

Protein structure databases with new web services for structural biology and biomedical research

The Protein Data Bank Japan (PDBj) curates, edits and distributes protein structural data as a member of the worldwide Protein Data Bank (wwPDB) and currently processes ~25–30% of all deposited data in the world. Structural information is enhanced by the addition of biological and biochemical functional data as well as experimental details extracted from the literature and other databases. Several applications have been developed at PDBj for structural biology and biomedical studies: (i) a Java-based molecular graphics viewer, jV; (ii) display of electron density maps for the evaluation of structure quality; (iii) an extensive database of molecular surfaces for functional sites, eF-site, as well as a search service for similar molecular surfaces, eF-seek; (iv) identification of sequence and structural neighbors; (v) a graphical user interface to all known protein folds with links to the above applications, Protein Globe. Recent examples are shown that highlight the utility of these tools in recognizing remote homologies between pairs of protein structures and in assigning putative biochemical functions to newly determined targets from structural genomics projects.

for more

Tuesday, April 15, 2008

Position open dbSNP Curator

The Single Nucleotide Polymorphisms database (dbSNP) serves as a central repository for both single base nucleotide substitutions and short deletion and insertion polymorphisms. Computercraft seeks a biologist with significant knowledge in life sciences to help curate dbSNP records, process submission, and perform data analyst tasks. Candidates should also have the ability to rapidly develop applications to process data into database and to generate reports.

The individual will work onsite at the National Institutes of Health (NIH) in Bethesda, MD. Our scientists work with genomic experts at NIH's National Center for Biotechnology Information (NCBI) in the National Library of Medicine (NLM) to create and enhance a suite of databases and tools available to researchers worldwide.

Requirements:
• PhD or M.S. in molecular biology, bioinformatics, or highly related field
• Linux/UNIX experience
• Relational database and SQL experience
• Programming experience (Perl, Python, or C++)
• Excellent verbal and written communication skills as well as organizational skills are essential
• Ability to work with a team in a production environment and communicate well with NCBI biologists and programmers as well as external scientist
• Strong interest in contributing to the development of public database resources
• Ideal candidate will have experience with dbSNP, genotyping, and QC processing

Other Desirable Experience:
• XML/XSLT
• Javascript and AJAX
• Web development

Additional information about the Single Nucleotide Polymorphisms is available at: http://www.ncbi.nlm.nih.gov/projects/SNP/


To apply for this position or learn about other Computercraft job opportunities, please visit the Careers section of our website: http://www.computercraft-usa.com/

Monday, March 31, 2008

CAS numbers are not public domain, are they?

"Work created before the existence of copyright and patent laws also form part of the public domain. The Bible and the inventions of Archimedes are in the public domain. However, copyright may exist in translations or new formulations of this work." [Wikipedia]
As posted by Tony is the Chemical Abstract Service (CAS) discouraging using their CAS services for assigning correct CAS numbers to structures for any third party database. Wikipedia is a source of structures, which is public domain due to its GNU FDL. Still, this does not imply that any translation of structures, e.g. CAS numbers, are in the public domain, too. Honestly, this raises a serious problem for curating CAS numbers on Wikipedia and this raises indeed the question, if they should not be dropped from Wikipedia, and any other information source, at all? Is it not better having no information, than having wrong information?

A CAS number is for me only one certain translation of a chemical structure. In this case, the only source and creator for CAS numbers is the American Chemical Society. CAS claims that their services can not be used for curating other data sources. Does this also mean that people can not use CAS numbers from publications?
"The public domain can also be defined in contrast to trademarks. Names, logos, and other identifying marks used in commerce can be restricted as proprietary trademarks for a single business to use. Trademarks can be maintained indefinitely, but they can also lapse through disuse, negligence, or widespread misuse, and enter the public domain. It is possible, however, for a lapsed trademark to become proprietary again, leaving the public domain." [Wikipedia]
And does this also mean that scientists are basically not allowed publishing CAS numbers and structures in scientific publications? How can CAS numbers then be used at all, if we can not store this information?

Many questions, and who will answer them? And, if they got answered what is the way forward for getting curated structures within the public domain? I would say (again, this time in nicer words): 'get organized scientists worldwide!' If CAS can do it, we can do it? It may take longer to get the party started, but if we do not start it will never happen.
"An ideal collaborative resource would be designed for large-scale data mining, contain curated historical data, and have data standards and deposition tools that could constantly bring in data from the published literature. ... In other words, the party might take longer to get started than hoped for, but it should be worth the wait." [M. Baker, DOI 10.1038/nrd2148]

Source: CAS numbers are not public domain, are they?