Makes sense to have data open so it can be used to mine for solutions and alternatives. Best practices in making this data findable and easily usable is also very useful.
Microsoft Throws Weight Behind Open Data Movement
Financial Times
Richard Waters
Microsoft has announced its support for the open data movement, urging governments and companies worldwide to share more data to ensure “digital power” is not concentrated in the U.S., China, and a handful of major technology companies. Microsoft pledged to make some of its own data available more widely, while developing standardized tools and legal frameworks to help others do the same. Microsoft president Brad Smith said the latest artificial intelligence innovations have raised the stakes, giving rise to "a looming data divide" that threatens to leave behind countries and companies with less access to data. Microsoft’s Jennifer Yokohama said the company is working on a shareable “live repository of best practices and resources,” along with “proof of concept initiatives to demonstrate how we can do open data better to really solve key societal challenges.” ... '
Showing posts with label Open Data. Show all posts
Showing posts with label Open Data. Show all posts
Wednesday, April 22, 2020
Saturday, March 07, 2020
Case for Open Data in Fight Against COVID-19
The case for open data for AI in the fight against COVID-19
Posted by ajit jaokar in DSC ....
COVID-19
2019 Novel Coronavirus COVID-19 (2019-nCoV) Data Repository by Johns Hopkins CSSE
This is the data repository for the 2019 Novel Coronavirus Visual Dashboard operated by the Johns Hopkins University Center for Systems Science and Engineering (JHU CSSE). Also, Supported by ESRI Living Atlas Team and the Johns Hopkins University Applied Physics Lab (JHU APL). ....
Note also the comments ...
Posted by ajit jaokar in DSC ....
COVID-19
2019 Novel Coronavirus COVID-19 (2019-nCoV) Data Repository by Johns Hopkins CSSE
This is the data repository for the 2019 Novel Coronavirus Visual Dashboard operated by the Johns Hopkins University Center for Systems Science and Engineering (JHU CSSE). Also, Supported by ESRI Living Atlas Team and the Johns Hopkins University Applied Physics Lab (JHU APL). ....
Note also the comments ...
Wednesday, October 16, 2019
Open Data vs Public Data
Short piece in DSC. Useful for my current work. Not sure if the definition is always applied this way, so you have to cautious in using the tag as others apply it.
Is There a Difference Between Open Data and Public Data?
Posted by Lewis Wynne-Jones
Yep. And it’s a big one.
There is a general consensus that when we talk about open data we are referring to any piece of data or content that is free to access, use, reuse, and redistribute. Due to the way most governments have rolled out their open data portals, however, it would be easy to assume that the data available on these sites is the only data that’s available for public consumption. This isn’t true.
Although data sets that receive a governmental stamp of openness receive a lot more publicity, they actually only represent a fraction of the public data that exists on the web. So what’s the difference between “public” data and “open” data?
What is open data? ... "
Is There a Difference Between Open Data and Public Data?
Posted by Lewis Wynne-Jones
Yep. And it’s a big one.
There is a general consensus that when we talk about open data we are referring to any piece of data or content that is free to access, use, reuse, and redistribute. Due to the way most governments have rolled out their open data portals, however, it would be easy to assume that the data available on these sites is the only data that’s available for public consumption. This isn’t true.
Although data sets that receive a governmental stamp of openness receive a lot more publicity, they actually only represent a fraction of the public data that exists on the web. So what’s the difference between “public” data and “open” data?
What is open data? ... "
Tuesday, March 05, 2019
Google Talks Finding, Sharing Open Data
Good talk on the issues, perhaps not enough on protecting the data, defining its ownership responsibly.
Doing our part to share open data responsibly
Daphne Luong Director, Software Engineering
Charina Chou Global Policy Lead, Emerging Technologies
This past weekend marked Open Data Day, an annual celebration of making data freely available to everyone. Communities around the world organized events, and we’re taking a moment here at Google to share our own perspective on the importance of open data. More accessible data can meaningfully help people and organizations, and we’re doing our part by opening datasets, providing access to APIs and aggregated product data, and developing tools to make data more accessible and useful.
Responsibly opening datasets
Sharing datasets is increasingly important as more people adopt machine learning through open frameworks like TensorFlow. We’ve released over 50 open datasets for other developers and researchers to use. These include YouTube 8M, a corpus of annotated videos used externally for video understanding; the HDR+ Burst Photography dataset, which helps others experiment with the technology that powers Pixel features like Portrait Mode; and Open Images, along with the Open Images Extended dataset which increases photo diversity.
Just because data is open doesn’t mean it will be useful, however. First, a dataset needs to be cleaned so that any insights developed from it are based on well-structured and accurate examples. Cleaning a large dataset is no small feat; before opening up our own, we spend hundreds of hours standardizing data and validating quality. Second, a dataset should be shared in a machine-readable format that’s easy for others to use, such as JSON rather than PDF. Finally, consider whether the dataset is representative of the intended content. Even if data is usable and representative of some situations, it may not be appropriate for every application. For instance, if a dataset contains mostly North American animal images, it may help you classify a deer, but not a giraffe. Tools like Facets can help you analyze the makeup of a dataset and evaluate the best ways to put it to use. We’re also working to build more representative datasets through interfaces like the Crowdsource application. To guide others’ use of your own dataset, consider publishing a data card which denotes authorship, composition and suggested use cases (here’s an example from our Open Images Extended release).
Making data findable and useful ... '
Doing our part to share open data responsibly
Daphne Luong Director, Software Engineering
Charina Chou Global Policy Lead, Emerging Technologies
This past weekend marked Open Data Day, an annual celebration of making data freely available to everyone. Communities around the world organized events, and we’re taking a moment here at Google to share our own perspective on the importance of open data. More accessible data can meaningfully help people and organizations, and we’re doing our part by opening datasets, providing access to APIs and aggregated product data, and developing tools to make data more accessible and useful.
Responsibly opening datasets
Sharing datasets is increasingly important as more people adopt machine learning through open frameworks like TensorFlow. We’ve released over 50 open datasets for other developers and researchers to use. These include YouTube 8M, a corpus of annotated videos used externally for video understanding; the HDR+ Burst Photography dataset, which helps others experiment with the technology that powers Pixel features like Portrait Mode; and Open Images, along with the Open Images Extended dataset which increases photo diversity.
Just because data is open doesn’t mean it will be useful, however. First, a dataset needs to be cleaned so that any insights developed from it are based on well-structured and accurate examples. Cleaning a large dataset is no small feat; before opening up our own, we spend hundreds of hours standardizing data and validating quality. Second, a dataset should be shared in a machine-readable format that’s easy for others to use, such as JSON rather than PDF. Finally, consider whether the dataset is representative of the intended content. Even if data is usable and representative of some situations, it may not be appropriate for every application. For instance, if a dataset contains mostly North American animal images, it may help you classify a deer, but not a giraffe. Tools like Facets can help you analyze the makeup of a dataset and evaluate the best ways to put it to use. We’re also working to build more representative datasets through interfaces like the Crowdsource application. To guide others’ use of your own dataset, consider publishing a data card which denotes authorship, composition and suggested use cases (here’s an example from our Open Images Extended release).
Making data findable and useful ... '
Friday, October 05, 2018
Free Open Source Training Data
Always looking for good open source data, especially for testing models. William Vorhies of of DSC talks this. Sources and links. Good examples:
Lots of Free Open Source Datasets to Make Your AI Better Posted by William Vorhies
Summary: There are several approaches to reducing the cost of training data for AI, one of which is to get it for free. Here are some excellent sources.
Recently we wrote that training data (not just data in general) is the new oil. It’s the difficulty and expense of acquiring labeled training data that causes many deep learning projects to be abandoned.
It also matters a great deal just how good you want your new deep learning app to be. A 2016 study by Goodfellow, Bengio and Courville concluded you could get ‘acceptable’ performance with about 5,000 labeled examples per category BUT it would take 10 Million labeled examples per category to “match or exceed human performance”.
There are a number of technologies coming up through research now that promise more accurate auto labeling to make creating training data less costly and time consuming. Snorkel from the Stanford Dawn Project is one we covered recently. This area is getting a lot of research attention. ... "
Lots of Free Open Source Datasets to Make Your AI Better Posted by William Vorhies
Summary: There are several approaches to reducing the cost of training data for AI, one of which is to get it for free. Here are some excellent sources.
Recently we wrote that training data (not just data in general) is the new oil. It’s the difficulty and expense of acquiring labeled training data that causes many deep learning projects to be abandoned.
It also matters a great deal just how good you want your new deep learning app to be. A 2016 study by Goodfellow, Bengio and Courville concluded you could get ‘acceptable’ performance with about 5,000 labeled examples per category BUT it would take 10 Million labeled examples per category to “match or exceed human performance”.
There are a number of technologies coming up through research now that promise more accurate auto labeling to make creating training data less costly and time consuming. Snorkel from the Stanford Dawn Project is one we covered recently. This area is getting a lot of research attention. ... "
Thursday, September 06, 2018
Google Announces Dataset Search
Looks useful, would have been useful in a number of cases. Will be interesting to see how well the search works, note the mention that it is similar to Google Scholar. Also an invitation to publish your own data indexed in this way. Could such a system be used to usefully scope what data exists and does not? Determine what a new data asset might be worth?
Making it Easier to Discover Datasets By Natasha Noy Research Scientist, Google AI
In today's world, scientists in many disciplines and a growing number of journalists live and breathe data. There are many thousands of data repositories on the web, providing access to millions of datasets; and local and national governments around the world publish their data as well. To enable easy access to this data, we launched Dataset Search, so that scientists, data journalists, data geeks, or anyone else can find the data required for their work and their stories, or simply to satisfy their intellectual curiosity.
There's a sea of open research data available on the web, but it can be time-consuming to sift through those sites to get at the data -- and it's not always presented in an easy-to-parse format. Google hopes it can make that information more accessible to scientists, journalists and plain old data junkies with its new Dataset Search feature. The tool provides more direct access to data presented in an open standard that makes it clear who created the info, how it was collected and how you're allowed to use it. You could not only track down climate data for a report, but make sure that it's relevant and legal to use.
Similar to how Google Scholar works, Dataset Search lets you find datasets wherever they’re hosted, whether it’s a publisher's site, a digital library, or an author's personal web page. To create Dataset search, we developed guidelines for dataset providers to describe their data in a way that Google (and other search engines) can better understand the content of their pages. These guidelines include salient information about datasets: who created the dataset, when it was published, how the data was collected, what the terms are for using the data, etc. We then collect and link this information, analyze where different versions of the same dataset might be, and find publications that may be describing or discussing the dataset. Our approach is based on an open standard for describing this information (schema.org) and anybody who publishes data can describe their dataset this way. We encourage dataset providers, large and small, to adopt this common standard so that all datasets are part of this robust ecosystem. ... "
Also in Engadget.
Making it Easier to Discover Datasets By Natasha Noy Research Scientist, Google AI
In today's world, scientists in many disciplines and a growing number of journalists live and breathe data. There are many thousands of data repositories on the web, providing access to millions of datasets; and local and national governments around the world publish their data as well. To enable easy access to this data, we launched Dataset Search, so that scientists, data journalists, data geeks, or anyone else can find the data required for their work and their stories, or simply to satisfy their intellectual curiosity.
There's a sea of open research data available on the web, but it can be time-consuming to sift through those sites to get at the data -- and it's not always presented in an easy-to-parse format. Google hopes it can make that information more accessible to scientists, journalists and plain old data junkies with its new Dataset Search feature. The tool provides more direct access to data presented in an open standard that makes it clear who created the info, how it was collected and how you're allowed to use it. You could not only track down climate data for a report, but make sure that it's relevant and legal to use.
Similar to how Google Scholar works, Dataset Search lets you find datasets wherever they’re hosted, whether it’s a publisher's site, a digital library, or an author's personal web page. To create Dataset search, we developed guidelines for dataset providers to describe their data in a way that Google (and other search engines) can better understand the content of their pages. These guidelines include salient information about datasets: who created the dataset, when it was published, how the data was collected, what the terms are for using the data, etc. We then collect and link this information, analyze where different versions of the same dataset might be, and find publications that may be describing or discussing the dataset. Our approach is based on an open standard for describing this information (schema.org) and anybody who publishes data can describe their dataset this way. We encourage dataset providers, large and small, to adopt this common standard so that all datasets are part of this robust ecosystem. ... "
Also in Engadget.
Wednesday, July 25, 2018
ISSIP Board Meeting
From the ISSIP Board meeting today See ISSIP.org
The International Society of Service Innovation Professionals, ISSIP (pronounced iZip), is a 501 (C) (3) professional association co-founded by IBM, Cisco, HP and several Universities with a mission to promote Service Innovation for our interconnected world. Our purpose is to help institutions and individuals to grow and be successful in our global service economy.
Service innovations improve the quality-of-life of individuals and the wealth of institutions, from businesses to nations that are increasingly dominated by service revenues and economics. Advances in information technology and policy support the rapid scaling of new service innovations in health, education, government, finance, hospitality, retail, communications, transportation, energy, utilities; even in advanced agricultural and manufacturing systems viewed as socio-technical systems, in which community-oriented recycling behaviors improve the economics, sustainability, and resilience of these human-serving systems. ...
---------------
ISSIP Open Datasets
Welcome to the ISSIP Open Data sets page! Here you will find a compiled list of hundreds of useful sites that have data sets to use for AI and Machine Learning. ...
Will follow with more information about our group. Join us.
The International Society of Service Innovation Professionals, ISSIP (pronounced iZip), is a 501 (C) (3) professional association co-founded by IBM, Cisco, HP and several Universities with a mission to promote Service Innovation for our interconnected world. Our purpose is to help institutions and individuals to grow and be successful in our global service economy.
Service innovations improve the quality-of-life of individuals and the wealth of institutions, from businesses to nations that are increasingly dominated by service revenues and economics. Advances in information technology and policy support the rapid scaling of new service innovations in health, education, government, finance, hospitality, retail, communications, transportation, energy, utilities; even in advanced agricultural and manufacturing systems viewed as socio-technical systems, in which community-oriented recycling behaviors improve the economics, sustainability, and resilience of these human-serving systems. ...
---------------
ISSIP Open Datasets
Welcome to the ISSIP Open Data sets page! Here you will find a compiled list of hundreds of useful sites that have data sets to use for AI and Machine Learning. ...
Will follow with more information about our group. Join us.
Wednesday, July 18, 2018
Microsoft Releases all US Building Footprints
Reported in Flowingdata, fascinating dataset. Architectural and building industry studies? An Exmple of open data
Details in Microsoft Github.
" ... This dataset contains 124,885,597 computer generated building footprints in all 50 US states. This data is freely available for download and use.
License
This data is licensed by Microsoft under the Open Data Commons Open Database License (ODbL)
FAQ
What the data include:
Approximately 125 million building footprint polygon geometries in all 50 US States in GeoJSON format. .... "
Details in Microsoft Github.
" ... This dataset contains 124,885,597 computer generated building footprints in all 50 US states. This data is freely available for download and use.
License
This data is licensed by Microsoft under the Open Data Commons Open Database License (ODbL)
FAQ
What the data include:
Approximately 125 million building footprint polygon geometries in all 50 US States in GeoJSON format. .... "
Monday, July 16, 2018
Microsoft Open Data
Brought to my attention: Microsoft Open Data. Their blog post about it.
A collection of free datasets from Microsoft Research to advance state-of-the-art research in areas such as natural language processing, computer vision, and domain specific sciences. Download or copy directly to a cloud-based Data Science Virtual Machine for a seamless development experience. ... "
Potentially very useful. https://msropendata.com/
A collection of free datasets from Microsoft Research to advance state-of-the-art research in areas such as natural language processing, computer vision, and domain specific sciences. Download or copy directly to a cloud-based Data Science Virtual Machine for a seamless development experience. ... "
Potentially very useful. https://msropendata.com/
Wednesday, May 16, 2018
Talk: Democratizing AI Through Open Data
Cognitive Systems Institute Talks:
17 May 2018: 10:30 AM, ET Access Instructions below.
Talk by: Michael Henretty
Title: “Finding our Common Voice: Democratizing AI through Open Data” Mozilla
Abstract:
Michael Henretty is an engineer and open innovation strategist working for Mozilla from Berlin. He is currently leading the Common Voice project, Mozilla's initiative to crowdsource a public dataset of human voices to be used in open speech technology. In a former life, Michael was a video game developer and designer.
The future of the Web is built on voice: voice recognition tools, speech directed commands, as well as civic movements. If we want to work towards a more inclusive, people enabled, and empowered future for the Internet as a global public resource, we need to act now. The challenge is complex, it starts with the training data used for speech algorithms, is connected to technical challenges for compatible code and systems, and finds its expression in continuously changing political contexts....
Slides now in place.
-------------------
Join the meetings by pointing your web browser to: https://zoom.us/j/7371462221 ; Callin: (415) 762-9988 or (646) 568-7788 Meeting id 7371462221 ; International Numbers: https://zoom.us/zoomconference.
Slides and Recording will be placed here: http://cognitive-science.info/community/weekly-update/
Join the CSIG LinkedIn Group to get reminders about talks and discuss them. Use twitter: #CSIGNews & #OpenTechAI
Replays before Dec 2015: Dial 877.471.6587 or 402.970.2667 and enter the call’s Replay ID when prompted for a program ID number. The Replay ID is listed in the Recording column of each date.
--------------------
Subscribe to:
Posts (Atom)