Showing posts with label curation. Show all posts
Showing posts with label curation. Show all posts

Monday, March 12, 2012

Database Curator@EBI for InterPro database

(EMBL-EBI seeking to recruit an enthusiastic Scientific Database Curator to join the InterPro team at the The European Bioinformatics Institute (EMBL-EBI) located on the Wellcome Trust Genome Campus near Cambridge in the UK.
The post-holder will work as part of a small team maintaining and curating the InterPro database. Their responsibilities will include updating existing InterPro entries, integrating new predictive signatures and adding annotation such as concise, literature-referenced abstracts. Data in InterPro needs to be of a consistently high quality and so potential candidates should have a good attention to detail and a thorough attitude to their work. We believe that understanding our users' needs and providing for them is critically important, and so the curator may be required to attend conferences, workshops or training events in order to meet with users, hear their ideas and expectations, and teach them about InterPro.
The EBI, part of the European Molecular Biology Laboratory (EMBL), provides cutting-edge research, services and training in the field of bioinformatics and is home to world-class resources such as UniProtKB and InterPro. InterPro generates and houses data predicting the functional classification of protein sequences and the presence of protein domains and sites. This information is widely used by genome sequencing projects and is disseminated to a large, global biological research community through web-based databases and software tools.
Requirements: 
The ideal applicant should hold at minimum a BSc. in a Biological Science, preferably with a strong background (e.g. 3 years post-graduate experience) in proteomics, molecular biology, biochemistry, cell biology and/or a related field.
Past work must either include work in a laboratory or in a biological database environment. A good understanding of proteomics and protein evolution would be advantageous, as would a high standard of scientific writing, with experience writing specifically for the web. Familiarity with the use of software tools for nucleotide/amino acid sequence analysis would be an advantage. The role is part of a tightly-knit team, so an ability to communicate ideas and openly discuss potential approaches will be very important. However, applicants should also be able to take their own initiative and work autonomously.
The candidate may be required to describe their work to a wider audience, including end-users of the database, so good presentation skills and an ability to write training materials and teach are a necessity; some of their work may also be published in scientific journals.
No computer programming skills are necessary but a proficiency at using Microsoft Office and a willingness to learn how to use new software tools are a must.
About Our Organization: 
EMBL is an inclusive, equal opportunity employer offering attractive conditions and benefits appropriate to an international research organisation.
Please note that appointments on fixed term contracts can be renewed, depending on circumstances at the time of the review.
Note that special visa requirements apply to employees from non EU countries working at EMBL-EBI in the UK. The period of work does not qualify for the Highly Skilled Migrants Programme.

Thursday, May 6, 2010

Opportunity: Microarray & Next-generation Sequence Curator for Gene Expression Omnibus (GEO) curation team

Computercraft seeks a highly motivated Molecular Biologist to join the Gene Expression Omnibus (GEO) curation team onsite at the National Institutes of Health (NIH) in Bethesda, MD. GEO is the largest fully public repository for functional genomic data, primarily microarray and next-generation sequence datasets. More information on GEO can be found on the web site http://www.ncbi.nlm.nih.gov/geo.

RESPONSIBILITIES:
We are currently looking for someone with a background in molecular biology, genomics, or biomedicine that is capable of working with large datasets. This person will be a member of the GEO curation team, helping to review and process incoming data submissions. The successful candidate will work at NIH's National Center for Biotechnology Information (NCBI) in the National Library of Medicine (NLM).

The successful candidate will perform the following tasks:
- Review and evaluate data submissions for structural integrity and content.
- Communicate extensively with researchers, resolving issues relating to submission procedures, formats, content, and site navigation.
- Utilize UNIX C shell commands and scripts to assemble, edit, and upload large data files to the database.
- Perform advanced database curation and assembly of comparable datasets to reflect biological variables and experimental design
- Test
, develop, and troubleshoot new query, data display, and analysis features.

REQUIREMENTS:
- This challenging position requires a Ph.D. or M.Sc in molecular biology or related field.
- Excellent general computer skills, including familiarity with spreadsheets, are required.
- Excellent written/verbal communication skills are an absolute requirement.

PREFERENCES:
- Practical experience with microarrays or high-throughput sequencing is highly desirable but not required to apply.
- Experience with LINUX/UNIX is highly desired.

SALARY & BENEFITS:
Computercraft offers a competitive salary and an excellent benefits package including PPO health insurance with 100% company paid premiums, 401K program with matching, paid time off and holiday pay, life insurance, flexible spending and disability coverage. We offer an excellent work life balance with a standard 40 hour work week and the chance to work alongside accomplished scientists at NIH/NCBI.

HOW TO APPLY:
To apply for this position or learn about other Computercraft job opportunities, please visit the Careers section of our website: http://www.computercraft-usa.com/.

Computercraft is an equal opportunity employer.

Tuesday, August 11, 2009

Curated databases and data curation

"There does appear to be a distinction between the way curation is used in the bio-sciences, and elsewhere. In particular, the term "curated database" tends to mean a manually constructed database that links literature to data, curated by experts who provide authority (eg see the Wikipedia definition of Biocurator). The earliest mention of the term "curated database" I can find is in the abstract (and only in the abstract) of Larsen et al (1993)." Chris Rusbridge Digital Curation Blog
Since these database are hand curated by experts (manually curated), they always promise a accuracy & quality better than uncurated or NLP based databases. While NLP based databases follow a automated curation provide quick updates and tend to be large in terms of the quantum of data. While they may trade off in accuracy due to their automated curation process.

While platforms like XTractor Premium follow a unique approach by trying to adopt the best of both worlds. A hybrid approach of a first level of NLP which promises a quick turnaround and large quantum of data followed by expert curation to preserve and maintain the high quality promised by expert curated database. XTractor Premium is a specialized biomedical text mining platform with semantically enriched search (Semantic search) and analytics to that enable discovery, knowledge sharing, analysis and modelling of published biomedical facts.
















Monday, October 27, 2008

Freely mining PUBMED for your drug discovery needs everyday

Biomedical Data mining happens to be a long-standing problem in scientific research. Scientists are constantly in search of newer and innovative means to mine biomedical data.

Be a part of the XTractor community. XTractor is the first of its kind - Literature alert service, that provides manually curated and annotated sentences for the Keywords of user preference. XTractor maps the extracted entities (genes, processes, drugs, diseases etc) to multiple ontologies and enables customized report generation. With XTractor the sentences are categorized into biological significant relationships and it also provides the user with the ability to create his own database for a set of Key terms. Also the user could change the Keywords of preference from time to time, with changing research needs. The categorized sentences could then be tagged and shared across multiple users. Thus XTractor proves to be a platform for getting real-time highly accurate data along with the ability to Share and collaborate.

Sign up it's free, and takes less than a minute. Just click here:www.xtractor.in.












Wednesday, September 17, 2008

The future of biocuration

Blogging about curation in the past several issues have been discussed but here is something very critical.
To thrive, the field that links biologists and their data urgently needs structure, recognition and support.
Biocuration, the activity of organizing, representing and making biological information accessible to both humans and computers, has become an essential part of biological discovery and biomedical research. But curation increasingly lags behind data generation in funding, development and recognition.

Three urgent actions to advance this key field. First, authors, journals and curators should immediately begin to work together to facilitate the exchange of data between journal publications and databases. Second, in the next five years, curators, researchers and university administrations should develop an accepted recognition structure to facilitate community-based curation efforts. Third, curators, researchers, academic institutions and funding agencies should, in the next ten years, increase the visibility and support of scientific curation as a professional career.

Failure to address these three issues will cause the available curated data to lag farther behind current biological knowledge. Researchers will observe an increasing occurrence of obvious gaps in knowledge. As these gaps expand, resources will become less effective for generating and testing hypotheses, and the usefulness of curated data will be seriously compromised. When all the data produced or published are curated to a high standard and made accessible as soon as they become available, biological research will be conducted in a manner that is quite unlike the way it is done now.

Researchers will be able to process massive amounts of complex data much more quickly. They will garner insight about the areas of their interest rapidly with the help of inference programs. Digesting information and generating hypotheses at the computer screen will be so much faster that researchers will get back to the bench quickly for more experiments. Experiments will be designed with more insight; this increased specificity will cause an exponential growth in knowledge, much as we are experiencing exponential growth in data today.

Also read this

Do you want to know more?

Be a part of the XTractor community. XTractor is the first of its kind - Literature alert service, that provides manually curated and annotated sentences for the Keywords of user preference. XTractor maps the extracted entities (genes, processes, drugs, diseases etc) to multiple ontologies and enables customized report generation. With XTractor the sentences are categorized into biological significant relationships and it also provides the user with the ability to create his own database for a set of Key terms. Also the user could change the Keywords of preference from time to time, with changing research needs. The categorized sentences could then be tagged and shared across multiple users. Thus XTractor proves to be a platform for getting real-time highly accurate data along with the ability to Share and collaborate.

Sign up it's free, and takes less than a minute. Just click here:www.xtractor.in.












Wednesday, August 20, 2008

XTractor crosses 150 user registration in 3 weeks!!!!


Be a part of the XTractor community. XTractor is the first of its kind - Literature alert service, that provides manually curated and annotated sentences for the Keywords of user preference. XTractor maps the extracted entities (genes, processes, drugs, diseases etc) to multiple ontologies and enables customized report generation. With XTractor the sentences are categorized into biological significant relationships and it also provides the user with the ability to create his own database for a set of Key terms. Also the user could change the Keywords of preference from time to time, with changing research needs. The categorized sentences could then be tagged and shared across multiple users. Thus XTractor proves to be a platform for getting real-time highly accurate data along with the ability to Share and collaborate.

Sign up it's free, and takes less than a minute. Just click here:www.xtractor.in.












Sunday, July 27, 2008

Molecular Connections perceives inorganic growth path

Molecular Connections, a life sciences informatics major, is now looking at acquiring small and medium-sized companies as part of its inorganic growth strategy.

"The acquisition will depend on good technology fit and domain expertise that could plug into our products and services,'' Jignesh Bhate, CEO, Molecular Connections Pvt Ltd told Pharmabiz.

"We are aggressively scouting for companies both in India and abroad. If we lay our hands on an international company, we will have access to their customers. In the case of an Indian company, we are essentially looking at a synergy in content and database,'' Bhate said.

To facilitate these acquisitions, Molecular Connections has identified three options to raise funds: internal accruals, foreign loans and venture funds. The mode of raising capital will depend on the size of the deal and the company is in talks with foreign banks. Around $2 million was raised from Barings Private Equity in 2006 for expansion.

Early this month, Molecular Connections was adjudged as the most promising company of the year by CNBC-ICICI for the Emerging India Awards out of more than 3.5 lakh entries.

The company, with sound financials, was also recognized as one of the fastest growing companies in the Technology Fast 50 India 2007 survey conducted by Deloitte Touche Tohmatsu, Asia Pacific. Molecular Connections was ranked 8th in India in the overall category.

The six-year-old company's product portfolio of informatics solutions focuses on protein interaction mechanisms of biological processes and also Biomarkers. Its maiden offering NetPro, the largest database globally on protein-protein has been upgraded with pharmacokinetics model. This is a market leader with repeat subscriptions from GSK, Aventis, Merck, Becton & Dickenson's bio-pharma unit and Galapagos NV, a drug discovery company.

The second product, CliPro's new version has been supplemented with biomarker data for clinical data. A recent product, Receptome is a comprehensive database of Functional Ligand-Receptor Complexes. Another new product, 'Xtractor' is a literature alert service, giving manually curated sentences of keywords as given by scientists.

All products are targeted at international customers because investments in informatics are increasing globally and the sector is getting its due importance. India pharma-biotech space will take a few years to utilize informatics' solutions offered from Molecular Connections, stated Bhate.

With collaborations helping to open a new dimension in research, Molecular Connections has partnered with the Bangalore-based Connexious Life Sciences for pathway work. Recently, it has teamed up with the Indian Institute of Science for a department of Biotechnology grant. The company's CYP database will be utilized by Prof P Kondaiah from the Department of Molecular Reproduction, Development & Genetics to customize drugs for the Indian population.

The global informatics market is valued at $6-8 billion. Of these four areas comprising validation & target identification, clinical development including toxicity studies, lead optimization studies, part of combinatorial chemistry is estimated at $6-8billion. Right now, Molecular Connections has the expertise in validation & target identification and is honing its skills in the area of ADME/TOX side and combinatorial chemistry.

Be a part of the XTractor community. XTractor is the first of its kind - Literature alert service, that provides manually curated and annotated sentences for the Keywords of user preference. XTractor maps the extracted entities (genes, processes, drugs, diseases etc) to multiple ontologies and enables customized report generation. With XTractor the sentences are categorized into biological significant relationships and it also provides the user with the ability to create his own database for a set of Key terms. Also the user could change the Keywords of preference from time to time, with changing research needs. The categorized sentences could then be tagged and shared across multiple users. Thus XTractor proves to be a platform for getting real-time highly accurate data along with the ability to Share and collaborate.

Sign up it's free, and takes less than a minute. Just click here:www.xtractor.in.










Wednesday, June 18, 2008

XTRACTOR™ Data Mining Simplified

The first of its kind - SCIENTIFIC LITERATURE alert service which also provides manually annotated sentences for the keywords of YOUR preference
  • Highly accurate, manually annotated sentences for given keywords.
  • Daily scientific literature updates at your desktop along with extracted facts - manually curated.
  • Provision to change keywords with your changing research preferences.
  • Annotated sentences and abstracts get stored in your profile, as and when they get updated in PubMed.
  • Classify and create your own datasets of annotated facts.
  • Enhanced experiences of reading & analyzing literature.
  • Access your profile / datasets from anytime, anywhere.
  • Discover and Create newer relations from scientific facts classified by the XTractor™ Community.
  • Tag your favorite abstracts and share them across other users.
  • Much more faster and an Absolutely Free Service

PRODUCT HIGHLIGHTS

Abstract Summarization
The XTractor™ system would provide extracted relations in addition to identifying the abstracts of subscriber's interest. Our experience suggests time required for summarization of an abstract is the greatest as the user needs to understand the context, analyze the content and make sense of the relation followed by extraction. So not only the abstracts would be prioritized and delivered to the subscriber but also sentences, which are extracted and processed manually.

Categorization
Categorization involves tagging the extracted sentences to most popular ontology(s) in biology and chemical spaces. XTractor™ would ensure highly accurate (manual curation) in all its annotation efforts be it genes, processes or drug names, all mapped to their relevant ontology(s). So you need not work again on reclassification of sentences or facts to accurate ontology(s).

Topic Tracking
XTractor™ provides updates to the subscribers on a daily basis based on topic tracking in PubMed. The service allows the user to Key in the Keywords of interest and notifies them via email when new articles get published in PubMed. Unlike the other free NLP utilities that are available, Molecular Connections would be using its skilled scientists to validate the data.

Do you want to know more?


Thursday, April 24, 2008

Structured Digital Abstracts - Easier Literature Searching

Already blogging on similar lines and the subject of Bring in data from the published literature in my earlier blog Smart tools to track, analyze and visualize research , here is another interesting experiment on similar lines:

The experiment centres on Structured Digital Abstracts (SDA). SDA are extensions of the normal journal article abstracts that describe the relationship between two biological entities, mentioning the method used to study the relationship. Each sentence is preceded by one or more identifiers pointing to the corresponding database entries that contain the full details of the interaction e.g. protein A interacts with protein B, by method X.

The aim of SDA is to assist data entry, text mining and literature searching by extracting the salient data from the article into simple sentences using a defined structure and controlled vocabularies.

Gianni Cesareni, Editor of FEBS Letters explains:

Many articles in biological journals describe relationships between entities (genes, proteins, etc.) yet this information cannot be efficiently used because of difficulties in retrieving from text. Databases capture this valuable information and organize it in a structured format ready for automatic analysis. The experiment of using SDAs will facilitate database entry and improve disclosure, to the benefit of authors and readers.


Tuesday, April 15, 2008

Position open dbSNP Curator

The Single Nucleotide Polymorphisms database (dbSNP) serves as a central repository for both single base nucleotide substitutions and short deletion and insertion polymorphisms. Computercraft seeks a biologist with significant knowledge in life sciences to help curate dbSNP records, process submission, and perform data analyst tasks. Candidates should also have the ability to rapidly develop applications to process data into database and to generate reports.

The individual will work onsite at the National Institutes of Health (NIH) in Bethesda, MD. Our scientists work with genomic experts at NIH's National Center for Biotechnology Information (NCBI) in the National Library of Medicine (NLM) to create and enhance a suite of databases and tools available to researchers worldwide.

Requirements:
• PhD or M.S. in molecular biology, bioinformatics, or highly related field
• Linux/UNIX experience
• Relational database and SQL experience
• Programming experience (Perl, Python, or C++)
• Excellent verbal and written communication skills as well as organizational skills are essential
• Ability to work with a team in a production environment and communicate well with NCBI biologists and programmers as well as external scientist
• Strong interest in contributing to the development of public database resources
• Ideal candidate will have experience with dbSNP, genotyping, and QC processing

Other Desirable Experience:
• XML/XSLT
• Javascript and AJAX
• Web development

Additional information about the Single Nucleotide Polymorphisms is available at: http://www.ncbi.nlm.nih.gov/projects/SNP/


To apply for this position or learn about other Computercraft job opportunities, please visit the Careers section of our website: http://www.computercraft-usa.com/

Monday, March 31, 2008

CAS numbers are not public domain, are they?

"Work created before the existence of copyright and patent laws also form part of the public domain. The Bible and the inventions of Archimedes are in the public domain. However, copyright may exist in translations or new formulations of this work." [Wikipedia]
As posted by Tony is the Chemical Abstract Service (CAS) discouraging using their CAS services for assigning correct CAS numbers to structures for any third party database. Wikipedia is a source of structures, which is public domain due to its GNU FDL. Still, this does not imply that any translation of structures, e.g. CAS numbers, are in the public domain, too. Honestly, this raises a serious problem for curating CAS numbers on Wikipedia and this raises indeed the question, if they should not be dropped from Wikipedia, and any other information source, at all? Is it not better having no information, than having wrong information?

A CAS number is for me only one certain translation of a chemical structure. In this case, the only source and creator for CAS numbers is the American Chemical Society. CAS claims that their services can not be used for curating other data sources. Does this also mean that people can not use CAS numbers from publications?
"The public domain can also be defined in contrast to trademarks. Names, logos, and other identifying marks used in commerce can be restricted as proprietary trademarks for a single business to use. Trademarks can be maintained indefinitely, but they can also lapse through disuse, negligence, or widespread misuse, and enter the public domain. It is possible, however, for a lapsed trademark to become proprietary again, leaving the public domain." [Wikipedia]
And does this also mean that scientists are basically not allowed publishing CAS numbers and structures in scientific publications? How can CAS numbers then be used at all, if we can not store this information?

Many questions, and who will answer them? And, if they got answered what is the way forward for getting curated structures within the public domain? I would say (again, this time in nicer words): 'get organized scientists worldwide!' If CAS can do it, we can do it? It may take longer to get the party started, but if we do not start it will never happen.
"An ideal collaborative resource would be designed for large-scale data mining, contain curated historical data, and have data standards and deposition tools that could constantly bring in data from the published literature. ... In other words, the party might take longer to get started than hoped for, but it should be worth the wait." [M. Baker, DOI 10.1038/nrd2148]

Source: CAS numbers are not public domain, are they?

Thursday, March 20, 2008

Smart tools to track, analyze and visualize research

Yes i am talking about the development in the field of literature search and scientific publication search. Looking beyond NLP some of the advancements and service models based on NLP and other advanced methods. Today to the user's delight there are quite a few companies and organization that provide services that absolutely make an scientist life easy!

These services revolve around customized alert services for selected search terms, specific research area of interest even more so pertaining to a list of custom key-word and search terms that the service providers allow the user to choose form. The scientist can simply subscribe to such service, sit back and wait for daily updates. The updates are provided on a day to day basis at his desktop.

He gets abstract summarization, full text summarization so that he need not actually go through the ordeal of reading through pages of scientific literature. Some once else does the job for him and just pick up the right things that he would look for in such papers. It is a very interactive service where on a daily basis the user can keep track of the quality of the service and set it right so that he can make sure that he just gets the right things he ever wants and nothing else other than that.

Over and above the service is provided using cutting edge technology, attractive web interface. Such interface come in all colour with floating dockable modules which is easy to use and some smart tools too. These smart tools enable the user to track, analyze, visualize and develop relationships among blocks of data that is being provided as "lego blocks".

"Knowledge is power" but knowledge is money too big money these days!

As an estimate companies providing such services are looking at a market size of $770 million worldwide which spans across verticals like bio-cheminformatics, knowledge management and content database. While the estimated market size for computational chemistry and knowledge management for life science vertical is approximately $400 million.


Monday, February 4, 2008

Life Science and Informatics

What is this?
is this a new industry?
or a old wine in a new bottle?

may be the later is true, bioinformatics is more popularly known by this name in the industry circle these days!

Well Life Sciences and Informatics can be any thing form anything form computational biology, all omes and omics, core bioinformatics to curation and literature mining, database creation, in the area of biology, chemistry , bio-chem space.

Cheminformatics, insilico chemistry and insilico biology how best to describe the host of activities involved that companies do to serve their clients and earn revenues for their firms...

There are number of companies in India and bangalore is the forefront as a major bio-cluster with 20 to 30 companies operation in this sphere.

there are companies with the sole focus as informatics as business
to name a few :

While IT companies which are also in to Life Sciences as one of their business verticals
  • TCS
  • HCL (not yet in Bangalore)
  • Wipro etc.
now how good are these companies doing?
how good are they in terms of the international markets and how profitable is their business?
what do they do?
their clients?

These are some interesting things that could be discussed in this blog page...

Links: