Welcome!

Apache Authors: Pat Romanski, Liz McMillan, Elizabeth White, Christopher Harrold, Janakiram MSV

Related Topics: @BigDataExpo, Artificial Intelligence, @CloudExpo, @ThingsExpo

@BigDataExpo: Blog Feed Post

Demystifying #DataScience | @CloudExpo #BigData #AI #ArtificialIntelligence

Data science is about identifying those variables and metrics that might be better predictors of performance

[Opening Scene]: Billy Dean is pacing the office. He’s struggling to keep his delivery trucks at full capacity and on the road. Random breakdowns, unexpected employee absences, and unscheduled truck maintenance are impacting bookings, revenues and ultimately customer satisfaction. He keeps hearing from his business customers how they are leveraging data science to improve their business operations. Billy Dean starts to wonder if data science can help him. As he contemplates what data science can do for him, he slowly drifts off to sleep, and visions of Data Science starts dancing in his head…

[Poof! Suddenly Wizard Wei appears]: Hi, I’m your data science wizard to help alleviate your data science concerns. I don’t understand why folks try to make the data science discussion complicated. Let’s start simple with a simple definition of data science:

Data science is about identifying those variables and metrics that might be better predictors of performance

The key to a successful analytical model is having a robust set of variables against which to test for their predictive capabilities. And the key to having a robust set of variables from which to test is to get the business users engaged early in the process.

[A confused Billy Dean]: Okay, but I’m still confused. I mean, how does this really apply to my business?

[A patient Wizard Wei]: Well, let’s say that you are trying to predict which of your routes are likely to have under-capacity loads so that you can combine loads. In order to identify those variables that might be better predictors of under-capacity routes, you might ask your business users:

What data might you want to have in order to predict under-capacity routes?

The business users are likely to come up with a wide variety of variables, including:

Customer name Ship to location Customer industry
Building permits Customer tenure Change in customer size
Customer stock price Customer D&B rating Types of products hauled
Time of year Seasonality/Holidays Day of week
Traffic Weather Local Events
Distance from distribution center Open headcount on Indeed.com Tenure of logistics manager

The Data Science team will then gather these variables, perform some data transformations and enrichment, and then look for variables and combinations of variables that yield the best predictive results regarding under-capacity routes (see Figure 1).

Figure 1: Data Science Process

Role of Artificial Intelligence
[A less confuse Billy Dean]:
Ah, I think I understand, but what about all this talk about artificial intelligence? From some of these commercials on TV, it appears that robots with artificial intelligence will be ruling the world. Can you say Skynet?

[A still patient Wizard Wei]: Ah, that’s just marketing. Artificial intelligence is just one of many different tools in the predictive analytics kit bag of a data scientist. But artificial intelligence – while embracing some very sophisticated mathematical, data enrichment and computing techniques – is really pretty straightforward. All artificial intelligence is trying to do is to find and quantify relationships between variables buried in large data sets (see Figure 2).

Figure 2: Understanding Artificial Intelligence

[An inquisitive Billy Dean]: Okay, I’m starting to get it, but there seems to be some many
different analytic and predictive algorithms from which to choose. How does the business user know where to start?

[A growing frustrated Wizard Wei]: Ah, that’s the secret to the process. Business users don’t need to know which algorithms to use; they need to be able to identify those variables that might be better predictors of performance. It is up to the data science team to determine which variables are the most appropriate by testing the different algorithms.

Data Mining, Machine Learning and Artificial Intelligence (including areas such as cognitive computing, statistics, neural networks, text analytics, video analytics, etc.) are all members of the broader category of data science tools. Our data scientist team has experts in each of these areas, though no one data scientist is an expert at all of them (in spite of what they tell me). The different data science tools are used in different scenarios for different needs. Think of one of your mechanics. They have a large toolbox full of different tools. They determine what tools to use to fix a truck based upon the problem they are trying to solve. That’s exactly what a data scientist is doing, just with a different toolbox of algorithms.

No single algorithm is best over whole domain; so different algorithms are needed to cover different domains. Often combinations of algorithms are used in order to achieve the best results. To be honest, it’s like a giant jigsaw puzzle with the data science team constantly testing different combinations of metrics, data enrichment and algorithms until they find the combination that yields the best results.

[An enlightened Billy Dean]: I think I’ve finally got it. All of these different algorithms and techniques are just trying to help predict what is likely to happen so that I can make better operational and customer issues. But what’s the realm of what’s possible with data and analytics; I mean, how effective can my organization become at leveraging data and analytics to power my business?

[A proud Wizard Wei]: Great question, and the heart of the big data and data science conversation. Figure 3 shows how you could use these different data science tools to progress up the Big Data Business Model Maturity Index; to transition from running your business on Descriptive analytics that tell you what happened (Monitoring stage) to Predictive analytics that tell you what is likely to happen (Insights stage) to Prescriptive analytics that tell you what they should do (Optimization stage).

Figure 3: Leveraging Artificial Intelligence to drive Business Value

In the end, the data and the analytics are only useful if they help you optimize key operational processes, reduce compliance and security risks, uncover new revenue opportunities and create a more compelling, more prescriptive customer engagement. In the end, data and analytics are all about your business.

[A satisfied Billy Dean]: That’s great Wizard Wei! Thanks for your help!

Now, what can you do about my taxes…

To learn more about “Demystifying Data Science”, come to my Dell EMC World session: “Demystifying Data Science: A Pragmatic Guide To Building Big Data Use Cases” See you there!!

The post Demystifying Data Science appeared first on InFocus Blog | Dell EMC Services.

More Stories By William Schmarzo

Bill Schmarzo, author of “Big Data: Understanding How Data Powers Big Business”, is responsible for setting the strategy and defining the Big Data service line offerings and capabilities for the EMC Global Services organization. As part of Bill’s CTO charter, he is responsible for working with organizations to help them identify where and how to start their big data journeys. He’s written several white papers, avid blogger and is a frequent speaker on the use of Big Data and advanced analytics to power organization’s key business initiatives. He also teaches the “Big Data MBA” at the University of San Francisco School of Management.

Bill has nearly three decades of experience in data warehousing, BI and analytics. Bill authored EMC’s Vision Workshop methodology that links an organization’s strategic business initiatives with their supporting data and analytic requirements, and co-authored with Ralph Kimball a series of articles on analytic applications. Bill has served on The Data Warehouse Institute’s faculty as the head of the analytic applications curriculum.

Previously, Bill was the Vice President of Advertiser Analytics at Yahoo and the Vice President of Analytic Applications at Business Objects.

@ThingsExpo Stories
"We are focused on SAP running in the clouds, to make this super easy because we believe in the tremendous value of those powerful worlds - SAP and the cloud," explained Frank Stienhans, CTO of Ocean9, Inc., in this SYS-CON.tv interview at 20th Cloud Expo, held June 6-8, 2017, at the Javits Center in New York City, NY.
"We've been engaging with a lot of customers including Panasonic, we've been involved with Cisco and now we're working with the U.S. government - the Department of Homeland Security," explained Peter Jung, Chief Product Officer at Pulzze Systems, in this SYS-CON.tv interview at @ThingsExpo, held June 6-8, 2017, at the Javits Center in New York City, NY.
In the enterprise today, connected IoT devices are everywhere – both inside and outside corporate environments. The need to identify, manage, control and secure a quickly growing web of connections and outside devices is making the already challenging task of security even more important, and onerous. In his session at @ThingsExpo, Rich Boyer, CISO and Chief Architect for Security at NTT i3, discussed new ways of thinking and the approaches needed to address the emerging challenges of security i...
What sort of WebRTC based applications can we expect to see over the next year and beyond? One way to predict development trends is to see what sorts of applications startups are building. In his session at @ThingsExpo, Arin Sime, founder of WebRTC.ventures, discussed the current and likely future trends in WebRTC application development based on real requests for custom applications from real customers, as well as other public sources of information.
The financial services market is one of the most data-driven industries in the world, yet it’s bogged down by legacy CPU technologies that simply can’t keep up with the task of querying and visualizing billions of records. In his session at 20th Cloud Expo, Karthik Lalithraj, a Principal Solutions Architect at Kinetica, discussed how the advent of advanced in-database analytics on the GPU makes it possible to run sophisticated data science workloads on the same database that is housing the rich...
SYS-CON Events announced today that Massive Networks will exhibit at SYS-CON's 21st International Cloud Expo®, which will take place on Oct 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. Massive Networks mission is simple. To help your business operate seamlessly with fast, reliable, and secure internet and network solutions. Improve your customer's experience with outstanding connections to your cloud.
DX World EXPO, LLC., a Lighthouse Point, Florida-based startup trade show producer and the creator of "DXWorldEXPO® - Digital Transformation Conference & Expo" has announced its executive management team. The team is headed by Levent Selamoglu, who has been named CEO. "Now is the time for a truly global DX event, to bring together the leading minds from the technology world in a conversation about Digital Transformation," he said in making the announcement.
Internet of @ThingsExpo, taking place October 31 - November 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA, is co-located with 21st Cloud Expo and will feature technical sessions from a rock star conference faculty and the leading industry players in the world. The Internet of Things (IoT) is the most profound change in personal and enterprise IT since the creation of the Worldwide Web more than 20 years ago. All major researchers estimate there will be tens of billions devic...
"The Striim platform is a full end-to-end streaming integration and analytics platform that is middleware that covers a lot of different use cases," explained Steve Wilkes, Founder and CTO at Striim, in this SYS-CON.tv interview at 20th Cloud Expo, held June 6-8, 2017, at the Javits Center in New York City, NY.
Everything run by electricity will eventually be connected to the Internet. Get ahead of the Internet of Things revolution and join Akvelon expert and IoT industry leader, Sergey Grebnov, in his session at @ThingsExpo, for an educational dive into the world of managing your home, workplace and all the devices they contain with the power of machine-based AI and intelligent Bot services for a completely streamlined experience.
With tough new regulations coming to Europe on data privacy in May 2018, Calligo will explain why in reality the effect is global and transforms how you consider critical data. EU GDPR fundamentally rewrites the rules for cloud, Big Data and IoT. In his session at 21st Cloud Expo, Adam Ryan, Vice President and General Manager EMEA at Calligo, will examine the regulations and provide insight on how it affects technology, challenges the established rules and will usher in new levels of diligence...
SYS-CON Events announced today that Calligo, an innovative cloud service provider offering mid-sized companies the highest levels of data privacy and security, has been named "Bronze Sponsor" of SYS-CON's 21st International Cloud Expo ®, which will take place on Oct 31 - Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. Calligo offers unparalleled application performance guarantees, commercial flexibility and a personalised support service from its globally located cloud plat...
SYS-CON Events announced today that DXWorldExpo has been named “Global Sponsor” of SYS-CON's 21st International Cloud Expo, which will take place on Oct 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. Digital Transformation is the key issue driving the global enterprise IT business. Digital Transformation is most prominent among Global 2000 enterprises and government institutions.
SYS-CON Events announced today that Datera, that offers a radically new data management architecture, has been named "Exhibitor" of SYS-CON's 21st International Cloud Expo ®, which will take place on Oct 31 - Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. Datera is transforming the traditional datacenter model through modern cloud simplicity. The technology industry is at another major inflection point. The rise of mobile, the Internet of Things, data storage and Big...
While the focus and objectives of IoT initiatives are many and diverse, they all share a few common attributes, and one of those is the network. Commonly, that network includes the Internet, over which there isn't any real control for performance and availability. Or is there? The current state of the art for Big Data analytics, as applied to network telemetry, offers new opportunities for improving and assuring operational integrity. In his session at @ThingsExpo, Jim Frey, Vice President of S...
"We provide IoT solutions. We provide the most compatible solutions for many applications. Our solutions are industry agnostic and also protocol agnostic," explained Richard Han, Head of Sales and Marketing and Engineering at Systena America, in this SYS-CON.tv interview at @ThingsExpo, held June 6-8, 2017, at the Javits Center in New York City, NY.
"DX encompasses the continuing technology revolution, and is addressing society's most important issues throughout the entire $78 trillion 21st-century global economy," said Roger Strukhoff, Conference Chair. "DX World Expo has organized these issues along 10 tracks with more than 150 of the world's top speakers coming to Istanbul to help change the world."
"MobiDev is a Ukraine-based software development company. We do mobile development, and we're specialists in that. But we do full stack software development for entrepreneurs, for emerging companies, and for enterprise ventures," explained Alan Winters, U.S. Head of Business Development at MobiDev, in this SYS-CON.tv interview at 20th Cloud Expo, held June 6-8, 2017, at the Javits Center in New York City, NY.
SYS-CON Events announced today that DXWorldExpo has been named “Global Sponsor” of SYS-CON's 21st International Cloud Expo, which will take place on Oct 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. Digital Transformation is the key issue driving the global enterprise IT business. Digital Transformation is most prominent among Global 2000 enterprises and government institutions.
In his opening keynote at 20th Cloud Expo, Michael Maximilien, Research Scientist, Architect, and Engineer at IBM, discussed the full potential of the cloud and social data requires artificial intelligence. By mixing Cloud Foundry and the rich set of Watson services, IBM's Bluemix is the best cloud operating system for enterprises today, providing rapid development and deployment of applications that can take advantage of the rich catalog of Watson services to help drive insights from the vast t...