Welcome!

Apache Authors: Pat Romanski, Liz McMillan, Elizabeth White, Christopher Harrold, Janakiram MSV

Related Topics: @BigDataExpo, @CloudExpo, Apache

@BigDataExpo: Blog Feed Post

The 'Thinking' Part of 'Thinking Like a Data Scientist' | @BigDataExpo #BigData

Unfortunately, people have a tendency to blindly trust a claim from any source that they deem credible

Imagine my surprise when reading the March 28, 2016 issue of BusinessWeek and stumbling across the article titled “Lies, Damned Lies, and More Statistics.” In the article, "BusinessWeek" warned readers to beware of “p-hacking,” which is the statistical practice of tweaking data in ways that generate low p-values but actually undermine the test (see p-value definition below). One of the results of “p-hacking” is that absurd results can be made to pass the p-value test, and important findings can be overlooked. For example…

A study from the Pennington Biomedical Research Center in Baton Rouge[1] followed 17,000 Canadians over 12 years and found that those who sat for most of the day were 54% more likely to die of heart attacks that those that didn’t.

54%!? Yikes, that’s a scary fact. Proof that sitting kills you by heart attack. As a person who spends a lot of time sitting behind a desk, or on an airplane, or at sporting events, this “54% more likely to die of heart attacks” fact is very concerning.  Can I cheat certain death by throwing out my current desk and buying one of those expensive “stand up” work desks? Sounds like a bargain.

But the BusinessWeek article concludes with the statement “… hold findings to a higher standard if they conflict with common sense.” Bottom-line: think!

Unfortunately, people have a tendency to blindly trust a claim from any source that they deem credible, even if it completely conflicts with their own experiences, or common sense.

It only takes a couple stats and lack of common sense to make a dangerous conclusion and claim it’s a fact. It’s harder to buy a gun in Illinois than most other states. Gun-related murder rates are higher in Illinois than most other states. So… we can conclude that stiffer gun laws cause murder. Right? No, we can’t conclude that from those stats.

But I started to think, and challenge the assumption that there is some sort of causality between sitting and heart attacks. Some questions that immediately popped to mind included:

  • Are there other variables, like lack of exercise or eating habits or age, which might be the cause of the heart attacks?
  • Was a control group used to test the validity of the study results?
  • Is there something about Canadians that makes them more susceptible to sitting and heart attacks?
  • Who sponsored this study? Maybe the manufacturer of these new expensive “stand up and work”-type desks?

One needs to be a bit skeptical when they hear these sorts of “factoids.” We should know better than to just believe these sorts of claims blindly.

Let’s use this to remind ourselves to think before jumping to conclusion. And this is a great opportunity to employ our “thinking like a data scientist” techniques to identify what other variables might contribute to this “54% more likely to die of heart attacks” observation. In particular, this is an opportunity to test the “By Analysis” to explore what other variables we might want to consider. To perform the “By Analysis,” let’s craft the statement against which we want to apply this technique:

“I want understand details on each of the study’s participants by…”

Here are some of the variables and metrics that we could test to see if they might be predictors of heart attacks:

Age

Gender

Health history

Critical health variables (e.g., weight, height, sedentary heart rate, active heart rate, BMI, LDL)

Cholesterol history

Exercise history

Historical exercise results

Family health history

Hours worked per day

Hours worked per week

Marital status

Divorced?

Number of dependents

Ages of dependents

Diet

 

Date of most recent vacation

Recent vacation location

Number of vacation days

Home location

Home weather

Work location

Work weather

Amount of airplane travel

Amount of car travel

Length of job commute

Job Stress

Job title

Life Stress

Years until retirement

Retirement readiness

Financial status

 

Upon further analysis, we start to see groupings of related variables and metrics around specific use cases including:

  • Lifestyle (age, gender, Body Mass Index, cholesterol levels, blood pressure, naps, hours of sleep, etc.)
  • Diet (calories, fat intake, sugar intake, alcohol, smoking, amount of fish, organic foods, etc.)
  • Exercise (frequency of exercise, recency of exercise, type of exercise, level of effort, heart rate, etc.)
  • Work (hours of work, days of work per week, stress levels, amount of airline travel, managing people, etc.)
  • Vacation (recency, frequency, days of vacation, location of vacation, vacation activities)
  • Environmental (number of sunny days, number of rainy days, number of grey days, range of temperatures, range of humidity, local traffic congestion, population density, amount of local green space, local entertainment, etc.)

And there are likely others that we would want to identify and test with respect to those variables and metrics ability to predict heart attacks.

The goal is to use these techniques to identify which metrics and variables are the best predictors of performance, and if you are doing that, you’re thinking like a data scientist!

If you are interested in learning more about this “Thinking Like A Data Scientist” process, join me at EMC World in Las Vegas (that should set off your stress alerts), on Monday, May 2, 12:00 PM – 1:00 PM. I promise not to stress you out!

Optional Reading: Hypothesis Testing, Null Hypothesis and p-values
Nerd Warning! I’m going to try to explain the p-value, but to do that I also need to explain the concepts of hypothesis testing and null hypothesis.

A hypothesis test is a statistical test to determine whether there is enough evidence in a sample of data to infer that a certain condition is true for the entire population. For example, the hypothesis to test is whether test group A had better cancer recovery results than test group B due to medication X. A hypothesis test examines two opposing hypotheses about a population: the null hypothesis and the alternative hypothesis.

The null hypothesis is the hypothesis that there is “no effect” or “no difference” between the test groups. The null hypothesis is a test that there is no significant difference between the test groups, and that any observed difference between the test groups is due to sampling or experimental error.

The alternative hypothesis (that there exists an effect or a difference between the test groups) is the hypothesis you want to be able to conclude is true.

You use a p-value to make the determination to reject the null hypothesis. If the p-value is less than or equal to the level of significance, which is a cut-off point that you define, then you can reject the null hypothesis. The smaller the p-value, the stronger the evidence to reject the null hypothesis (i.e., no significant difference between test groups) in favor of the alternative hypothesis (i.e., there is significant difference between test groups).

A common misconception is that statistical hypothesis tests are designed to select the more likely of two hypotheses. Instead, a hypothesis test only tests whether to reject the null hypothesis.

By the way, even if we “fail to reject” the null hypothesis, it does not mean the null hypothesis is true. That’s because a hypothesis test does not determine which hypothesis is true; it only assesses whether available evidence exists to reject the null hypothesis[2].

Confusing? Yea, that’s what I think as well.

The Minitab Blog was a great source for much of the above data (http://blog.minitab.com)

[1] Sources: “Cutting Daily Sitting Time to Under 3 Hours Might Extend Life by Two Years; Watching TV for Less Than 2 Hours a Day Might Add Extra 1.4 Years”, July 10, 2012 and “Sedentary behaviour and life expectancy in the USA: a cause-deleted life table analysis” BJ Open Accessible Medical Research

[2] Check out “Bewildering Things Statisticians Say: “Failure to Reject the Null Hypothesis” for more details on failing to reject the null hypothesis.

The post The “Thinking” Part of “Thinking Like A Data Scientist” appeared first on InFocus.

Read the original blog entry...

More Stories By William Schmarzo

Bill Schmarzo, author of “Big Data: Understanding How Data Powers Big Business”, is responsible for setting the strategy and defining the Big Data service line offerings and capabilities for the EMC Global Services organization. As part of Bill’s CTO charter, he is responsible for working with organizations to help them identify where and how to start their big data journeys. He’s written several white papers, avid blogger and is a frequent speaker on the use of Big Data and advanced analytics to power organization’s key business initiatives. He also teaches the “Big Data MBA” at the University of San Francisco School of Management.

Bill has nearly three decades of experience in data warehousing, BI and analytics. Bill authored EMC’s Vision Workshop methodology that links an organization’s strategic business initiatives with their supporting data and analytic requirements, and co-authored with Ralph Kimball a series of articles on analytic applications. Bill has served on The Data Warehouse Institute’s faculty as the head of the analytic applications curriculum.

Previously, Bill was the Vice President of Advertiser Analytics at Yahoo and the Vice President of Analytic Applications at Business Objects.

@ThingsExpo Stories
SYS-CON Events announced today that Secure Channels, a cybersecurity firm, will exhibit at SYS-CON's 21st International Cloud Expo®, which will take place on Oct 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. Secure Channels, Inc. offers several products and solutions to its many clients, helping them protect critical data from being compromised and access to computer networks from the unauthorized. The company develops comprehensive data encryption security strategie...
SYS-CON Events announced today that App2Cloud will exhibit at SYS-CON's 21st International Cloud Expo®, which will take place on Oct. 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. App2Cloud is an online Platform, specializing in migrating legacy applications to any Cloud Providers (AWS, Azure, Google Cloud).
WebRTC is the future of browser-to-browser communications, and continues to make inroads into the traditional, difficult, plug-in web communications world. The 6th WebRTC Summit continues our tradition of delivering the latest and greatest presentations within the world of WebRTC. Topics include voice calling, video chat, P2P file sharing, and use cases that have already leveraged the power and convenience of WebRTC.
Internet-of-Things discussions can end up either going down the consumer gadget rabbit hole or focused on the sort of data logging that industrial manufacturers have been doing forever. However, in fact, companies today are already using IoT data both to optimize their operational technology and to improve the experience of customer interactions in novel ways. In his session at @ThingsExpo, Gordon Haff, Red Hat Technology Evangelist, shared examples from a wide range of industries – including en...
Detecting internal user threats in the Big Data eco-system is challenging and cumbersome. Many organizations monitor internal usage of the Big Data eco-system using a set of alerts. This is not a scalable process given the increase in the number of alerts with the accelerating growth in data volume and user base. Organizations are increasingly leveraging machine learning to monitor only those data elements that are sensitive and critical, autonomously establish monitoring policies, and to detect...
To get the most out of their data, successful companies are not focusing on queries and data lakes, they are actively integrating analytics into their operations with a data-first application development approach. Real-time adjustments to improve revenues, reduce costs, or mitigate risk rely on applications that minimize latency on a variety of data sources. Jack Norris reviews best practices to show how companies develop, deploy, and dynamically update these applications and how this data-first...
Intelligent Automation is now one of the key business imperatives for CIOs and CISOs impacting all areas of business today. In his session at 21st Cloud Expo, Brian Boeggeman, VP Alliances & Partnerships at Ayehu, will talk about how business value is created and delivered through intelligent automation to today’s enterprises. The open ecosystem platform approach toward Intelligent Automation that Ayehu delivers to the market is core to enabling the creation of the self-driving enterprise.
"We're a cybersecurity firm that specializes in engineering security solutions both at the software and hardware level. Security cannot be an after-the-fact afterthought, which is what it's become," stated Richard Blech, Chief Executive Officer at Secure Channels, in this SYS-CON.tv interview at @ThingsExpo, held November 1-3, 2016, at the Santa Clara Convention Center in Santa Clara, CA.
Consumers increasingly expect their electronic "things" to be connected to smart phones, tablets and the Internet. When that thing happens to be a medical device, the risks and benefits of connectivity must be carefully weighed. Once the decision is made that connecting the device is beneficial, medical device manufacturers must design their products to maintain patient safety and prevent compromised personal health information in the face of cybersecurity threats. In his session at @ThingsExpo...
SYS-CON Events announced today that Massive Networks will exhibit at SYS-CON's 21st International Cloud Expo®, which will take place on Oct 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. Massive Networks mission is simple. To help your business operate seamlessly with fast, reliable, and secure internet and network solutions. Improve your customer's experience with outstanding connections to your cloud.
The question before companies today is not whether to become intelligent, it’s a question of how and how fast. The key is to adopt and deploy an intelligent application strategy while simultaneously preparing to scale that intelligence. In her session at 21st Cloud Expo, Sangeeta Chakraborty, Chief Customer Officer at Ayasdi, will provide a tactical framework to become a truly intelligent enterprise, including how to identify the right applications for AI, how to build a Center of Excellence to ...
From 2013, NTT Communications has been providing cPaaS service, SkyWay. Its customer’s expectations for leveraging WebRTC technology are not only typical real-time communication use cases such as Web conference, remote education, but also IoT use cases such as remote camera monitoring, smart-glass, and robotic. Because of this, NTT Communications has numerous IoT business use-cases that its customers are developing on top of PaaS. WebRTC will lead IoT businesses to be more innovative and address...
Everything run by electricity will eventually be connected to the Internet. Get ahead of the Internet of Things revolution and join Akvelon expert and IoT industry leader, Sergey Grebnov, in his session at @ThingsExpo, for an educational dive into the world of managing your home, workplace and all the devices they contain with the power of machine-based AI and intelligent Bot services for a completely streamlined experience.
Because IoT devices are deployed in mission-critical environments more than ever before, it’s increasingly imperative they be truly smart. IoT sensors simply stockpiling data isn’t useful. IoT must be artificially and naturally intelligent in order to provide more value In his session at @ThingsExpo, John Crupi, Vice President and Engineering System Architect at Greenwave Systems, will discuss how IoT artificial intelligence (AI) can be carried out via edge analytics and machine learning techn...
SYS-CON Events announced today that GrapeUp, the leading provider of rapid product development at the speed of business, will exhibit at SYS-CON's 21st International Cloud Expo®, which will take place October 31-November 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. Grape Up is a software company, specialized in cloud native application development and professional services related to Cloud Foundry PaaS. With five expert teams that operate in various sectors of the market acr...
SYS-CON Events announced today that Datera, that offers a radically new data management architecture, has been named "Exhibitor" of SYS-CON's 21st International Cloud Expo ®, which will take place on Oct 31 - Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. Datera is transforming the traditional datacenter model through modern cloud simplicity. The technology industry is at another major inflection point. The rise of mobile, the Internet of Things, data storage and Big...
In his opening keynote at 20th Cloud Expo, Michael Maximilien, Research Scientist, Architect, and Engineer at IBM, discussed the full potential of the cloud and social data requires artificial intelligence. By mixing Cloud Foundry and the rich set of Watson services, IBM's Bluemix is the best cloud operating system for enterprises today, providing rapid development and deployment of applications that can take advantage of the rich catalog of Watson services to help drive insights from the vast t...
SYS-CON Events announced today that CA Technologies has been named "Platinum Sponsor" of SYS-CON's 21st International Cloud Expo®, which will take place October 31-November 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. CA Technologies helps customers succeed in a future where every business - from apparel to energy - is being rewritten by software. From planning to development to management to security, CA creates software that fuels transformation for companies in the applic...
Recently, IoT seems emerging as a solution vehicle for data analytics on real-world scenarios from setting a room temperature setting to predicting a component failure of an aircraft. Compared with developing an application or deploying a cloud service, is an IoT solution unique? If so, how? How does a typical IoT solution architecture consist? And what are the essential components and how are they relevant to each other? How does the security play out? What are the best practices in formulating...
In his session at @ThingsExpo, Arvind Radhakrishnen discussed how IoT offers new business models in banking and financial services organizations with the capability to revolutionize products, payments, channels, business processes and asset management built on strong architectural foundation. The following topics were covered: How IoT stands to impact various business parameters including customer experience, cost and risk management within BFS organizations.