Had not seen this effort yet, worth a look.
Accelerating data-driven discoveries
Life science companies use Paradigm4’s unique database management system to uncover new insights into human health.
Zach Winn | MIT News Office
As technologies like single-cell genomic sequencing, enhanced biomedical imaging, and medical “internet of things” devices proliferate, key discoveries about human health are increasingly found within vast troves of complex life science and health data.
But drawing meaningful conclusions from that data is a difficult problem that can involve piecing together different data types and manipulating huge data sets in response to varying scientific inquiries. The problem is as much about computer science as it is about other areas of science. That’s where Paradigm4 comes in.
The company, founded by Marilyn Matz SM ’80 and Turing Award winner and MIT Professor Michael Stonebraker, helps pharmaceutical companies, research institutes, and biotech companies turn data into insights.
It accomplishes this with a computational database management system that’s built from the ground up to host the diverse, multifaceted data at the frontiers of life science research. That includes data from sources like national biobanks, clinical trials, the medical internet of things, human cell atlases, medical images, environmental factors, and multi-omics, a field that includes the study of genomes, microbiomes, metabolomes, and more.
On top of the system’s unique architecture, the company has also built data preparation, metadata management, and analytics tools to help users find the important patterns and correlations lurking within all those numbers.
In many instances, customers are exploring data sets the founders say are too large and complex to be represented effectively by traditional database management systems.
“We’re keen to enable scientists and data scientists to do things they couldn’t do before by making it easier for them to deal with large-scale computation and machine-learning on diverse data,” Matz says. “We’re helping scientists and bioinformaticists with collaborative, reproducible research to ask and answer hard questions faster.” .... " unique database management system to uncover new insights into human health. ... "
Showing posts with label Bioinformatics. Show all posts
Showing posts with label Bioinformatics. Show all posts
Sunday, April 05, 2020
Thursday, March 30, 2017
Streamlined Text Mining
A time ago we used relatively simplistic methods of context analysis to understand the meaning of gathered text. Now with focused methods and increasing amounts of text data, new results are now possible. In Datanami:
Biomedical Text Mining Tool Gets the Lead Out by George Leopold
Approximately 100 lines of Python code serve as the basis of a new predictive text-mining tool designed to accelerate the scanning online biomedical research papers for clues on everything from repurposing existing drugs to advancing stem cell treatment.
Coders from the Morgridge Institute for Research working in partnership with the University of Wisconsin at Madison reported on their “KinderMiner” algorithm during a bioinformatics conference in San Francisco this week. The researchers said the 100-line algorithm was “within hours” able to scan more than 30 million online papers to provide ranked and relevant associations based on key words and phrases.
“Most often, researchers are running manual Google searches and combing through millions of hits to find, for example, certain genes that are important to a biological process or disease,” explained Ron Stewart, associate director of bioinformatics at the Morgridge Institute. “It’s often based on hunches and intuition. We’re trying to automate and formalize that process.”
Alternative techniques require much data wrangling, added Finn Kuusisto, a postdoctoral researcher at the Morgridge Institute. “We write about 100 lines of Python code, and our users can be given answers that may significantly speed up their scientific process.”
Biomedical Text Mining Tool Gets the Lead Out by George Leopold
Approximately 100 lines of Python code serve as the basis of a new predictive text-mining tool designed to accelerate the scanning online biomedical research papers for clues on everything from repurposing existing drugs to advancing stem cell treatment.
Coders from the Morgridge Institute for Research working in partnership with the University of Wisconsin at Madison reported on their “KinderMiner” algorithm during a bioinformatics conference in San Francisco this week. The researchers said the 100-line algorithm was “within hours” able to scan more than 30 million online papers to provide ranked and relevant associations based on key words and phrases.
“Most often, researchers are running manual Google searches and combing through millions of hits to find, for example, certain genes that are important to a biological process or disease,” explained Ron Stewart, associate director of bioinformatics at the Morgridge Institute. “It’s often based on hunches and intuition. We’re trying to automate and formalize that process.”
Alternative techniques require much data wrangling, added Finn Kuusisto, a postdoctoral researcher at the Morgridge Institute. “We write about 100 lines of Python code, and our users can be given answers that may significantly speed up their scientific process.”
Saturday, June 13, 2015
Online Data Science Apprenticeship
Live and ready, just updated, DSC's online apprenticeship. Have scanned their work and it seems to be well done. Impressed by the work that DSC has done providing information about data science topics in general. Worth a look, Get started here.
" ... This is an ideal program for professionals with a quantitative background and some industry experience - in a nutshell, for anyone who understands our cheat sheet and can get started using it. This program is for self-learners. Completed projects, submitted to and reviewed by Dr. Vincent Granville, will be featured, published, and promoted on our network, reaching out to the largest audience of data science decision makers, peers and hiring managers. Examples of such projects can be found here and here.
Ideal for people who want to change career paths, consultants, people managing data scientists, or students starting an analytic degree. In short, the easy way to become a data scientist for educated self-learners.
Free from unnecessary advanced matrix algebra or obscure, confusing, misused stats developed before the era of computers (p-value, GLM, model-based predictive analytics). Instead focusing on building automated, black-box, simple, scalable, efficient, robust, high ROI solutions for many modern applications (API's, machine-to-machine communications and real time), leveraging distributed architectures as needed, and using a unified cross-disciplinary approach. Applications: IoT, digital (mobile) and web data, marketing, risk management, operations research, finance, bioinformatics, healthcare, environmental data science, engineering, business analytics, model-free predictive analytics, business optimization, big data and more. .... "
" ... This is an ideal program for professionals with a quantitative background and some industry experience - in a nutshell, for anyone who understands our cheat sheet and can get started using it. This program is for self-learners. Completed projects, submitted to and reviewed by Dr. Vincent Granville, will be featured, published, and promoted on our network, reaching out to the largest audience of data science decision makers, peers and hiring managers. Examples of such projects can be found here and here.
Ideal for people who want to change career paths, consultants, people managing data scientists, or students starting an analytic degree. In short, the easy way to become a data scientist for educated self-learners.
Free from unnecessary advanced matrix algebra or obscure, confusing, misused stats developed before the era of computers (p-value, GLM, model-based predictive analytics). Instead focusing on building automated, black-box, simple, scalable, efficient, robust, high ROI solutions for many modern applications (API's, machine-to-machine communications and real time), leveraging distributed architectures as needed, and using a unified cross-disciplinary approach. Applications: IoT, digital (mobile) and web data, marketing, risk management, operations research, finance, bioinformatics, healthcare, environmental data science, engineering, business analytics, model-free predictive analytics, business optimization, big data and more. .... "
Tuesday, November 02, 2010
Data Scope Project
In CACM: A quite interesting proposal. Don't quite understand the details involved as yet, but the analogy is an exciting thing. Understanding complex data often includes problems of scale ...how do we choose the proper granularity of the data to solve the decision problem at hand? Could give a new view into the scope of business intelligence ...
Data-Scope Computer to Allow Data Analysis That's Impossible Today
Imagine a tool that is a cross between a powerful electron microscope and the Hubble Space Telescope, allowing scientists from disciplines ranging from medicine and genetics to astrophysics, environmental science, oceanography and bioinformatics to examine and analyze enormous amounts of data from both "little picture" and "big picture" perspectives..... "
Data-Scope Computer to Allow Data Analysis That's Impossible Today
Imagine a tool that is a cross between a powerful electron microscope and the Hubble Space Telescope, allowing scientists from disciplines ranging from medicine and genetics to astrophysics, environmental science, oceanography and bioinformatics to examine and analyze enormous amounts of data from both "little picture" and "big picture" perspectives..... "
Subscribe to:
Posts (Atom)