/* ---- Google Analytics Code Below */
Showing posts with label Teradata. Show all posts
Showing posts with label Teradata. Show all posts

Saturday, October 01, 2022

Teradata's Data Lakehouse

 Do I need a Lakehouse? 

Data Infrastructure by TeraData

Last week Teradata offered its long-awaited response to the emergence of the data lakehouse. As VentureBeat’s George Lawton reported, Teradata has always differentiated itself by stretching the capabilities of analytics, first with massively parallel processing on its own specialized machines, and more recently, with software-defined appliances tuned for variations in workloads — from compute-intensive to IOPS (input/output operations per second)-intensive. And since the acquisition of Aster Data Systems over a decade ago, Teradata morphed from solving big analytics problems to solving any analytics problem with a diverse portfolio of analytic libraries stretching SQL to new areas such as path or graph analytics.

With the cloud, we’ve been waiting for when Teradata would fully exploit cloud object storage, which is the de facto data lake. So the dual announcements last week of VantageCloud Lake Edition and ClearScape Analytics were logical next steps on Teradata’s journey to the data lakehouse. Teradata is finally making cloud storage a first-class citizen and opening it up to its wide analytics portfolio.

Summit Cloud in a Web 3.0 World How Edge computing drives platforms that converge the needs of IT, Ops, and Developers_Landscape

But unlike Teradata’s previous moves to parallelized and polyglot analytics, where it led the field, this time with the lakehouse, it has company. The announcement might not have mentioned the lakehouse word, but that’s what it was all about. As we noted several months back, almost everyone in the data world including Oracle, Teradata, Cloudera, Talend, Google, HPE, Fivetran, AWS, Dremio and even Snowflake has felt compelled to respond to Databricks, which introduced the data lakehouse.

Teradata’s path to the data lakehouse

Nonetheless, Teradata approaches the data lakehouse with some unique twists and is all about optimization. Teradata’s secret sauce has always been about highly optimized compute, interconnects, storage and query engines, along with workload management designed to run compute resources up to 95% utilization. When commodity hardware got good enough, Teradata introduced IntelliFlex where performance and optimizations could be configured through software. The capability to optimize for hardware not-invented-here opened the door to Teradata optimizing for AWS, and down the road, the other hyperscalers.

MetaBeat will bring together thought leaders to give guidance on how metaverse technology will transform the way all industries communicate and do business on October 4 in San Francisco, CA.

Teradata introduced VantageCloud a year ago, and late last year ran a 1,000+ node benchmark that no other cloud analytics provider has so far matched. But this was for a more conventional data warehouse using customary block storage.

The complication in making the lakehouse happen was developing a table format for data sitting in cloud object storage. That allows all the niceties associated with data warehouses, such as ACID transactions, which are key to ensuring consistency of data, more granular security and access controls, and raw performance. Databricks fired the first shot with Delta Lake, and more recently, other providers from Snowflake to Cloudera and others have embraced Apache Iceberg, the common thread being that this is all based on open source technology. For Lake Edition, Teradata went its own way with its own data lake table format, which the company claims delivers superior performance compared to Delta and Iceberg.

The other side of the lakehouse coin is software. Aside from its SQL engine, which has been designed to handle large, complex queries that can join up to hundreds of tables, Teradata has a large portfolio of analytic libraries that run in-database. This has been one of Teradata’s best-kept secrets. Largely the legacy of the Aster Data acquisition over a decade ago, these analytics were specially tuned to exploit the underlying parallelism, and they went well beyond SQL, encompassing functions such as n-Path, graph, time series analysis, and machine learning, all accessed through SQL extensions.  ... '

Wednesday, March 23, 2022

Teradata Tests it's Cloud

Have no heard from this direction of late.

Teradata Puts New Cloud Architecture to the 1,000-Node Test  By Alex Woodie

Teradata says the recent 1,000-node test that it ran on AWS not only shows the scale that its new cloud architecture can achieve, but also demonstrates the analytic flexibility required by its new target market, the Global 10,000.

Teradata is a company on the move, and it’s moving both in terms of what it makes and whom it makes it for. It’s no longer developing tightly coupled, on-prem data warehouses for the largest corporations in the world. Instead, it’s creating de-coupled analytic software that runs anywhere–on prem, in the cloud, or in multiple clouds–and its market has widened considerably, from the Fortune 500 to the Global 10,000.

This background is important to understand why the Teradata Innovation Lab spent the time to take Teradata Vantage (the name of its cloud-based offering) for a 1,000-node spin. According to Teradata Chief Product Officer Hillary Ashton, the AWS cluster was more than twice as big as any production cluster run by a Teradata customer.

“We have some of the largest customers on-premises on the planet, so I think it’s a really great indication of where we’re heading in the future,” Ashton told Datanami. “Obviously it gives our large enterprise customers the comfort that our future is big enough for the largest enterprise workloads.”

The test, which took about four weeks to run and was announced one week ago, utilized 100 TB of data and simulated thousands of concurrent SQL queries submitted by more than 1,000 simulated users. The workload itself was a mix of quick-hitting operational queries that demanded fast response times, as well as more complex and longer-running decision support system (DSS) queries.

There are some details of the test missing, including the specific EC2 instance types that were used (Teradata says there were two types used), and the total cost of the system. The company says it will share more details in a white paper that will be published in a couple of weeks. But don’t expect pricing information, as this test was never intended to be a public price-performance benchmark. .....'

Saturday, June 15, 2019

History and Current State of Teradata

Have encountered Teradata a number of times as it related to retail analyics.  They turn 40 and encounter yet more competition.   A considerable piece on their state and future.

Teradata Turns 40, Takes Off Gloves, Readies for a Fight    By Alex Woodie in DataNami

The past, present, and future of Teradata collided yesterday at the company’s headquarters in San Diego, California, where 1,500 employees and guests – including four of the company’s original founders — gathered to celebrate the company’s improbable 40-year run. And facing tough competition in the analytics ring, Teradata’s CEO vows to go on the offensive.

The odds were stacked against Teradata from the very beginning, when a group of computer scientists gathered in Jack Shemer’s garage in Brentwood, California. It was the summer of 1979, and the future was not what you would call bright.

“All I remember is interest rates were astronomically high,” said Jerry Modes, one of Teradata’s co-founders, in a video that played on a large outdoor screen. “Carter was president….We decided to start a company right in the worst time there was.”

“Between Labor Day and Thanksgiving, the price of gold went from $350 to $850 dollars an ounce,” said Phillip Naches, another co-founder, in the video (which you can view here). “The prime rate went from 8% to 22%, and we knew that there was a big recession was coming on.”

“It was not an environment conducive to being a startup,” Shemer said in the video. “Being a startup, you know you’re going to live or you’re going to die.  You don’t want to fall on your sword or fall on your six shooter. You want to make sure you take a rifle shot and succeed.”

The founders were all technically adept, but they relied on Shemer, Teradata’s original CEO, to shake some money loose from the venture capitalists, which was not an easy task given the economy at the time.

“He could charm the birds out of the trees,” Teradata co-founder David Hartke said in the video. “He went around to many investors and talked them into investing with Teradata.”  .... "

Thursday, January 17, 2019

Doing AI with all of your Data

Been a while since I have read the Teradata blog, some good thoughts.    This is the challenge for many enterprises.  And their operational data today is still relational.

Using AI on All Your Data: Building AI Models with Relational Data
By Ben MacKenzie in the Teradata Blog

As I have written elsewhere, the most striking advances in AI in the last few years have been in computer vision, natural language processing, and reinforcement learning:   think of image classification, Google Translate, and AlphaGo.   What has been overlooked is that innovations in machine and deep learning is also quietly revolutionizing the analysis of tabular, or relational, data.  The less-hyped advances in the analysis of relational data will have far-reaching consequences for enterprises, since every enterprise has relational data.  Powerful new techniques for relational data will enable enterprises with the right technology partners to make better decisions faster, with all of their data, all of the time. However, there are hurdles to achieving a future state of pervasive data intelligence across all enterprise data.  In this blog, I will reflect on some of the many challenges particular to building models with relational data.

The most obvious difference is the data itself.  When you are building a computer vision or natural language model, you start with a static set of images associated with categories, or a static corpus of text paired with corresponding translations. Acquiring these data sets, especially the labeled data that is essential for training the models, can be very difficult.  However, the data acquisition challenges are fundamentally different from those involved in acquiring a relational data set.  Relational data is the lifeblood of an organization.  Like blood, it flows into and out of different parts of an organization, where it lives in diverse operational databases, typically using different data management schemes, and is subject to diverse security and privacy constraints. And yet, having a single, coherent, complete view of all the data across the organization is critical to the success of an AI effort.  Achieving such a single perspective on all enterprise data is a job for an enterprise data warehouse, whether it’s a product like Teradata Vantage, a well-designed and implemented Data Lake, or a logical data warehouse combining data warehouse and Data Lake.  A data warehouse is your unified  ... ' 

Wednesday, August 10, 2016

Scaling Data Science in R

Ran into exactly this problem recently.   Mostly we solved using R as an abstraction layer with sampled data.   There are a number of other solutions in the article:

You’ve got three options: Scaling up, scaling out, or using R as an abstraction layer.
By Federico Castanedo 

For more on this topic, Brian Kreeger and Roger Fried will be hosting a live webcast, Scalable Data Science with R, on August 16, 2016.

R is among the top five data science tools in use today according O’Reilly research; the latest kdnuggets survey puts it in first, and IEEE Spectrum ranks it as the fifth most popular programming language.

The latest Rexer Data Miner survey revealed that in the past eight years, there has been an three-fold increase in the number of respondents using R, and a seven-fold increase in the number of analysts/scientists who have said that R is their primary tool.

Despite its popularity, the main drawback of vanilla R is its inherently “single threaded” nature and its need to fit all the data being processed in RAM. But nowadays, data sets are typically in the range of GBs and they are growing quickly to TBs. In short, current growth in data volume and variety is demanding more efficient tools by data scientists   ... " 

Wednesday, March 30, 2016

Individualized Marketing

In the Teradata Blog:  Good view of the topic from a data perspective.   " ...  Make the Right Connection … Every Time ... So what is individualized marketing, and what does it mean for you? It means that you must ensure that every interaction you have with each customer is relevant to that individual, and offers must be easily adapted to his or her evolving needs.  ... " 

Wednesday, January 20, 2016

Analyzing the Internet of Things

Teradata's Bill Franks on the IOT.   Nice concise thoughts.  Like Big Data, the IOT's engine is analytics and then translating that to decison.  Not sure we need another acronymn (AoT), but I accept it as a reminder of the fundamental need for value from data.

" .... To date, a lot of effort has been put into creating sensors, deploying them, and generating masses of data. However, lagging behind that effort is the analysis of the data. As with any data, no value is driven without analysis and action. It would have been better if more thought was given to how to utilize the data generated prior to creating sensors that stream it out. Given that we are where we are, the best path forward is to begin to aggressively analyze the data of the IoT. This is what I, and others, have begun to call the Analytics of Things (AoT). ... "

Monday, December 07, 2015

Netflix, Analytics and What You Watch

In FutureStartup: Correspondent Bill Franks chief scientist at Teradata, talks about his analytics experiences with Netflix.    Further thoughts: " ... Teradata’s Franks said that if it wanted to, Netflix could almost be a stand-alone analytics firm.  ... “It is getting blurry out there as to what companies are in,” said Franks. “AT&T is not just a phone provider. It is providing TV, and then has its own data. Nike is now manufacturing high-tech electronics and housing data in its data center, not just doing knitted sportswear. Companies can commercialize the use of their data outside of their core business. It is a fascinating time, and a lot of companies will be morphing and twisting.” ... " 

Thursday, October 29, 2015

Teradata Listener

Analyzing real time streams of data,  from real systems, or from simulations or prototypes or models, is a useful concept.

Teradata doubles down on IoT with two new tools .... 
" ... The Internet of Things is a big part of what makes big data so big, and on Monday, analytics provider Teradata unleashed two new tools to help make sense of the vast troves of data it produces.

Teradata Listener and Teradata Aster Analytics on Hadoop help users "listen" to massive streams of IoT data in real time and then use analytics to find the distinctive underlying patterns.

Teradata Listener is self-service software for ingesting and distributing individual or multiple data streams from sources including sensors, telematics, mobile events, click streams, social media feeds and IT server logs. ... " 

Monday, September 28, 2015

Big Data Ducklings

In Teradata Mag:   " ... Ugly Duckling or Black Swan?  Big Data gives businesses the periheral vision to detect and respond to catastropic events before it is too late. ... " .  Thoughtful piece,  but I define black swans as those that cannot be completely predicted.  So this is more like risk management, which has its own literature.  If I build a portfolio of risks, and can define potential occurrences, then Black Swans are very low probability possibilities.  How low?  Completely unanticipated?   Events that I do not have lots of data about, are by definition not 'big' data.  Yes I know it is not all about 'big', which is part of the reason I do not like the term.  So lets just call it data analytics under unusual contexts.

Wednesday, July 22, 2015

Teradata Aster Videos on Advanced Analytics

Links to all videos  by John Thuma to demonstrate Teradata Aster analytics use.   Useful for training in this space.  He writes:  " ... Please see the following list of videos that I have created to document some of Asters 120+ Analytic Functions: ... " .   

 I have looked at a number of these,  which were short, as non technical as possible, and in general relate to real life applications.  Though the applications that they use may not match your industry context. Specifics of Aster use and training/testing models also included.   Nicely done. Useful for any analyst with some advanced analytics training.

Thursday, July 09, 2015

Discovery Analytics

Bill Franks of Teradata on Discovery analytics.  We used to call it  exploratory analytics, and every project started with some of that.  Which is why data visualization was so useful, even before we had so many tool options.   Good thoughts, he writes:  "  ...  I spend a lot of time these days talking with companies about the need for a formal approach to enabling what is often called “discovery analytics” or “exploratory analytics.” What I find is that many people have a fundamental misunderstanding of what discvery analytics is all about. There is one analogy that I have found to be effective in getting people to better understand the concept. In this blog, I’ll walk you through that analogy.  ... "      I agree.   Worth reading his complete argument.

How does this differ from the Watson Discovery Advisor, which I have written about before?  See the Watson Discovery Advisor tag.   The Discovery Advisor is more knowledge management than analytics oriented.  But it would make sense to combine the two for many applications.

Monday, July 06, 2015

Supply Chain Planning Garbage in Equals Garbage Out

Very interesting piece, which addresses a problem we often encountered, and was too often not addressed well enough.  Every enterprise should have a handle on cleansing and trends in data.  Read the whole thing for a solution view.

By Cheryl Wiebe,  Partner, Applied Analytics, Manufacturing at Teradata

Supply Chain Planning: Garbage In Garbage Out...
Lean manufacturing has been massively adopted over the last 10 years. Lean is driving attention to the notion of supply chain variability and accuracy of planning factors. Kanban, or pull style, says that the impetus for manufacturing activity (e.g., arrival of starting kit to begin assembly) is the arrival of the required materials at your station. This means if we have insufficient material to kit the line, the line goes down. 

What this means? Lean assumes your supply chain is infallible. This forces suppliers to plan JIT warehouses at the beginning of the line. A certain disk drive manufacturer (supplier A) had to supply 18,000 drives at the JIT warehouse nearby the customer's PC assembly plant, for example. If they dipped below 18k, the customer would not pull from supplier A, and would pull instead from supplier B, the competitor. This caused supplier A to ensure they always had 25k drives (excess inventory). A short on inventory from suppliers will bring down a Lean manufacturing line. The supply problem has now moved into the JIT warehouse. We have moved the problem upstream. ... " 

Saturday, June 27, 2015

Teradata Aster

An overvew of Teradata Aster and their community.  We visited Aster at a trip to Stanford in around 2010, before their acquisition by Teradata.  A number of integrated products and solutions.  I am now revisiting. Installed Aster Express.   Any inputs?    They write:

Looking for data-driven business insights?
With the integrated Teradata Aster Discovery Platform, organizations attain unmatched competitive advantage by making it faster and easier for a wider group of users to generate powerful, high impact business insights from big data. ... " 

Wednesday, May 20, 2015

Black Swans and Big Data

In TeraData Mag:  Good piece that looks at the inevitability of these events, and how Data can still help. Leads to the suggestion that predictive analytics will allow us to be prepared for classes of Black Swans.  Segmenting appropriate responses.  And then, obviously addressing them with the same patterns:   

" ....  A data-driven analysis or simulation designed to determine an organization’s ability to deal with a crisis situation can help it be ready for the day a black swan lands on its front steps. Krishna notes that this type of “stress testing” can gauge the level of readiness. “[Preparation] is really all about imagining the unimaginable, understanding what’s going to blow up ... [and] to be able to determine what corrective actions are needed,” he explains. “That is becoming very much an accepted approach ... certainly something that, from a regulatory standpoint, is becoming mandatory for a number of financial institutions.”

Businesses that identify a possible event early on and take evasive actions are usually in the best position to ride out the crisis. Krishna advises companies to carefully map out the exact steps they will need to take in various types of crisis situations. “Then document those actions so that if any of the expected scenarios occur, there will be no second-guessing,” he points out. “It’s simply a matter of executing what’s already been documented; executing that game plan, if you will.” ... ' 

Tuesday, May 12, 2015

Teradata Row and Column Innovation

Been working with Teradata since using RetailLink to leverage WalMart data.   Now examining the value of this new innovation.   Could it have been used with RetailLink?

" ... Teradata delivers World's most advanced hybrid row and column database ... Innovations allow customers to mix and match row and column technologies at will, without the trade-offs of purely columnar databases ... " 

Monday, April 20, 2015

Sunday, March 01, 2015

Big Data Strategy

Teradata Magazine on Big Data Strategy.   Nicely done, readily scannable information and graphics.   Useful statistics and detail.  A number of useful case studies.

Thursday, February 12, 2015

Big Data Apps from Teradata

Have just had cause to go back and look to see what Teradata has done in the Big Data space.    Interesting capabilities.  Taking a deeper dive. " ... To ease companies into realizing bankable big data benefits, Teradata has developed a collection of big data apps – pre-built templates that act as time-saving short cuts to value. Limited skill sets and complexity make it challenging for analytic professionals to rapidly and consistently derive actionable insights that can be easily operationalized.  Teradata is taking the lead in offering advanced analytic apps powered by Teradata Aster AppCenter to give sophisticated results from big data analytics. ...  "

More from Teradata on Big Data Apps.

Friday, November 07, 2014

Advanced Analytics in Consumer Goods: Podcast

Abstract: In Justin Honaman, managing partner for consumer goods for North American industry consulting at Teradata, is interviewed by Ron Powell, independent analyst and consultant focused big data, business intelligence and data warehousing. Ron and Justin discuss the why advanced analytics has become a priority for today’s consumer goods companies. Justin also talks about how consumer goods companies interact more directly with their consumers and why these companies now want to be data driven. He explains how the most successful CPG companies are the ones prioritizing data and information. Podcast.