Click here to close now.


Apache Authors: Craig Lowell, Jim Scott, Liz McMillan, AppDynamics Blog, Dana Gardner

News Feed Item

Cloudera Qualifies Data Scientists With New Certification Program

Hands-On Certification Prepares Data Scientists for Success With Real-World Data; Data Science Challenge Begins March 31

PALO ALTO, CA -- (Marketwired) -- 03/26/14 -- Cloudera, the leader in enterprise analytic data management powered by Apache Hadoop™, today announced the industry's first hands-on data science certification, called Cloudera Certified Professional: Data Scientist (CCP: DS). Comprised of a Data Science Essentials exam, a twice-annual Data Science Challenge, and several preparatory and enablement resources, Cloudera's data scientist certification program helps developers, analysts, statisticians, and engineers get experience with relevant big data tools and techniques and validate their abilities while helping prospective employers identify elite, highly skilled data scientists. The next Cloudera Data Science Challenge begins March 31, 2014.

Industry Faces Shortage of Qualified Data Scientists

Enterprises are increasingly storing massive amounts of data in Hadoop to streamline the path to actionable insights, develop advanced analytics models, and build big data tools that were previously unattainable for most organizations. As a result, the demand for data scientists is at an all-time high. Data scientists possess a rare combination of engineering capabilities, statistical skills, and subject matter expertise that is difficult to find. Job openings for data scientists far outpace the limited supply of these highly in-demand workers, and the skills gap is widening. The situation is complicated by the fact that there has historically not been a clearly established skill set or university degree that an individual could acquire to qualify as a data scientist. Companies seeking to hire their first data scientists often have little idea what credentials to look for in a candidate.

Cloudera Addresses Demand for Data Scientists through Training and Certification

As the global leader in Hadoop training and professional certification, Cloudera is addressing the widespread industry need for data scientists with its new CCP:DS certification. Designed and led by Cloudera's own elite team of data scientists, the CCP:DS program helps aspiring data scientists develop and prove out the skills they need to succeed with real-world enterprise data.

In addition to the certification exam, the program includes an optional three-day Introduction to Data Science course focused on teaching data professionals to build machine learning models and implement complex recommender systems with Hadoop as a platform using industry-standard tools like Python and Apache Mahout. Cloudera also offers a 60-question Data Science Essentials Practice Test for candidates to self-assess their exam-readiness, and a free Data Science Challenge Solution Kit consisting of a live data set, a step-by-step tutorial, and a detailed explanation of the processes required to arrive at the correct outcomes for real-world data science questions focused on classification, clustering, and collaborative filtering of web analytics.

Once candidates have passed the Data Science Essentials exam, they must then successfully complete a Cloudera Data Science Challenge, offered twice annually. By passing Cloudera's examination and live-data challenge, CCP:DS-credentialed individuals have demonstrated their ability to work with big data and build market-relevant data science models under real-world conditions at the very highest level. Cloudera Certified Professional: Data Scientist is the world's only certification that provides evidence of true experience and expertise developing a production-ready data science solution that is peer-evaluated for accuracy, scalability, and robustness.

Introducing the Data Science Challenge: Detecting Anomalies in Medicare Claims
Cloudera's second Data Science Challenge opens on March 31, 2014. Participants will have three months to complete the challenge. Designed by Cloudera's Director of Data Science, Sean Owen, the Data Science Challenge asks aspiring data scientists to detect possible errors and anomalies in Medicare claims using a massive set of anonymized healthcare data. Successful challengers will be able to answer questions, including:

  • Which medical procedures have the highest relative variance in cost?
  • Which three providers had the highest average amount claimed for the largest number of procedures?
  • Based on amount and type of procedures claimed, which three providers and regions are least like the others?
  • Identify 10,000 patients that seem most likely to need review for possible errors or anomalies. Describe some common features in these patients.

To learn more about the Data Science Challenge or to register, please visit:

Join us for a webinar about the current Data Science Challenge on April 10:

What Data Scientists Say about CCP:DS:
"The certification program that Cloudera has put together goes beyond the written test, including a challenge that is designed to assess the data scientist skills in much greater depth than could be achieved in a multiple choice questionnaire. From my perspective, this makes the exercise much more compelling, valuable, and meaningful than any other certification available today. You are actually solving problems through data analysis in a full simulation of situations data scientists face in the field."
- Luis Quintela, Samsung SDS, Cloudera Certified Professional: Data Scientist

"CCP:DS goes a long way towards removing ambiguity about who and what a data scientist is. Being associated with Cloudera earns instant respect, as well. Because the exam is based on real-world challenges and is fully vetted by some of the world's top experts, the certification does the hard work of pre-evaluating candidates against the multiple highly technical areas that would otherwise be difficult to qualify."
- David F. McCoy, confidential employer, Cloudera Certified Professional: Data Scientist

"I'm pumped to earn the CCP:DS credential! It holds true weight in the market because it replicates a real, sufficiently difficult big data scenario I would see on the job and requires a professional-level approach to solving problems. The exam captured all the relevant elements of data science and machine learning, and the challenge made the experience completely non-trivial."
- Stuart Horsman, Cloudera, Cloudera Certified Professional: Data Scientist

Learn More About Cloudera's Training and Professional Certification Programs
To learn more about Cloudera's comprehensive offering of big data training programs and professional certifications, including the new CCP: Data Scientist program, please visit:

About Cloudera
Cloudera is revolutionizing enterprise data management by offering the first unified Platform for Big Data, an enterprise data hub built on Apache Hadoop™. Cloudera offers enterprises one place to store, process and analyze all their data, empowering them to extend the value of existing investments while enabling fundamental new ways to derive value from their data. Only Cloudera offers everything needed on a journey to an enterprise data hub, including software for business critical data challenges such as storage, access, management, analysis, security and search. As the leading educator of Hadoop professionals, Cloudera has trained over 20,000 individuals worldwide. Over 900 partners and a seasoned professional services team help deliver greater time to value. Finally, only Cloudera provides proactive and predictive support to run an enterprise data hub with confidence. Leading organizations in every industry plus top public sector organizations globally run Cloudera in production.

Connect with Cloudera
Read our Vision blog:
Follow Cloudera on Twitter:
Follow Cloudera University on Twitter:
Visit us on Facebook:

Cloudera, Cloudera Platform for Big Data, Cloudera Enterprise Basic Edition, Cloudera Enterprise Flex Edition, Cloudera Enterprise Data Hub Edition and CDH are trademarks or registered trademarks of Cloudera in the United States and in jurisdictions throughout the world. All other company and product names may be trade names or trademarks of their respective owners.

Add to Digg Bookmark with Add to Newsvine

More Stories By Marketwired .

Copyright © 2009 Marketwired. All rights reserved. All the news releases provided by Marketwired are copyrighted. Any forms of copying other than an individual user's personal reference without express written permission is prohibited. Further distribution of these materials is strictly forbidden, including but not limited to, posting, emailing, faxing, archiving in a public database, redistributing via a computer network or in a printed form.

@ThingsExpo Stories
As more intelligent IoT applications shift into gear, they’re merging into the ever-increasing traffic flow of the Internet. It won’t be long before we experience bottlenecks, as IoT traffic peaks during rush hours. Organizations that are unprepared will find themselves by the side of the road unable to cross back into the fast lane. As billions of new devices begin to communicate and exchange data – will your infrastructure be scalable enough to handle this new interconnected world?
This week, the team assembled in NYC for @Cloud Expo 2015 and @ThingsExpo 2015. For the past four years, this has been a must-attend event for MetraTech. We were happy to once again join industry visionaries, colleagues, customers and even competitors to share and explore the ways in which the Internet of Things (IoT) will impact our industry. Over the course of the show, we discussed the types of challenges we will collectively need to solve to capitalize on the opportunity IoT presents.
SYS-CON Events announced today that Dyn, the worldwide leader in Internet Performance, will exhibit at SYS-CON's 17th International Cloud Expo®, which will take place on November 3-5, 2015, at the Santa Clara Convention Center in Santa Clara, CA. Dyn is a cloud-based Internet Performance company. Dyn helps companies monitor, control, and optimize online infrastructure for an exceptional end-user experience. Through a world-class network and unrivaled, objective intelligence into Internet conditions, Dyn ensures traffic gets delivered faster, safer, and more reliably than ever.
SYS-CON Events announced today that Sandy Carter, IBM General Manager Cloud Ecosystem and Developers, and a Social Business Evangelist, will keynote at the 17th International Cloud Expo®, which will take place on November 3–5, 2015, at the Santa Clara Convention Center in Santa Clara, CA.
SYS-CON Events announced today that Super Micro Computer, Inc., a global leader in high-performance, high-efficiency server, storage technology and green computing, will exhibit at the 17th International Cloud Expo®, which will take place on November 3–5, 2015, at the Santa Clara Convention Center in Santa Clara, CA. Supermicro (NASDAQ: SMCI), the leading innovator in high-performance, high-efficiency server technology is a premier provider of advanced server Building Block Solutions® for Data Center, Cloud Computing, Enterprise IT, Hadoop/Big Data, HPC and Embedded Systems worldwide. Supermi...
The Internet of Things (IoT) is growing rapidly by extending current technologies, products and networks. By 2020, Cisco estimates there will be 50 billion connected devices. Gartner has forecast revenues of over $300 billion, just to IoT suppliers. Now is the time to figure out how you’ll make money – not just create innovative products. With hundreds of new products and companies jumping into the IoT fray every month, there’s no shortage of innovation. Despite this, McKinsey/VisionMobile data shows "less than 10 percent of IoT developers are making enough to support a reasonably sized team....
The IoT market is on track to hit $7.1 trillion in 2020. The reality is that only a handful of companies are ready for this massive demand. There are a lot of barriers, paint points, traps, and hidden roadblocks. How can we deal with these issues and challenges? The paradigm has changed. Old-style ad-hoc trial-and-error ways will certainly lead you to the dead end. What is mandatory is an overarching and adaptive approach to effectively handle the rapid changes and exponential growth.
With major technology companies and startups seriously embracing IoT strategies, now is the perfect time to attend @ThingsExpo in Silicon Valley. Learn what is going on, contribute to the discussions, and ensure that your enterprise is as "IoT-Ready" as it can be! Internet of @ThingsExpo, taking place Nov 3-5, 2015, at the Santa Clara Convention Center in Santa Clara, CA, is co-located with 17th Cloud Expo and will feature technical sessions from a rock star conference faculty and the leading industry players in the world. The Internet of Things (IoT) is the most profound change in personal an...
The IoT is upon us, but today’s databases, built on 30-year-old math, require multiple platforms to create a single solution. Data demands of the IoT require Big Data systems that can handle ingest, transactions and analytics concurrently adapting to varied situations as they occur, with speed at scale. In his session at @ThingsExpo, Chad Jones, chief strategy officer at Deep Information Sciences, will look differently at IoT data so enterprises can fully leverage their IoT potential. He’ll share tips on how to speed up business initiatives, harness Big Data and remain one step ahead by apply...
There will be 20 billion IoT devices connected to the Internet soon. What if we could control these devices with our voice, mind, or gestures? What if we could teach these devices how to talk to each other? What if these devices could learn how to interact with us (and each other) to make our lives better? What if Jarvis was real? How can I gain these super powers? In his session at 17th Cloud Expo, Chris Matthieu, co-founder and CTO of Octoblu, will show you!
Developing software for the Internet of Things (IoT) comes with its own set of challenges. Security, privacy, and unified standards are a few key issues. In addition, each IoT product is comprised of at least three separate application components: the software embedded in the device, the backend big-data service, and the mobile application for the end user's controls. Each component is developed by a different team, using different technologies and practices, and deployed to a different stack/target - this makes the integration of these separate pipelines and the coordination of software upd...
As a company adopts a DevOps approach to software development, what are key things that both the Dev and Ops side of the business must keep in mind to ensure effective continuous delivery? In his session at DevOps Summit, Mark Hydar, Head of DevOps, Ericsson TV Platforms, will share best practices and provide helpful tips for Ops teams to adopt an open line of communication with the development side of the house to ensure success between the two sides.
Today air travel is a minefield of delays, hassles and customer disappointment. Airlines struggle to revitalize the experience. GE and M2Mi will demonstrate practical examples of how IoT solutions are helping airlines bring back personalization, reduce trip time and improve reliability. In their session at @ThingsExpo, Shyam Varan Nath, Principal Architect with GE, and Dr. Sarah Cooper, M2Mi's VP Business Development and Engineering, will explore the IoT cloud-based platform technologies driving this change including privacy controls, data transparency and integration of real time context w...
The Internet of Everything is re-shaping technology trends–moving away from “request/response” architecture to an “always-on” Streaming Web where data is in constant motion and secure, reliable communication is an absolute necessity. As more and more THINGS go online, the challenges that developers will need to address will only increase exponentially. In his session at @ThingsExpo, Todd Greene, Founder & CEO of PubNub, will explore the current state of IoT connectivity and review key trends and technology requirements that will drive the Internet of Things from hype to reality.
"Matrix is an ambitious open standard and implementation that's set up to break down the fragmentation problems that exist in IP messaging and VoIP communication," explained John Woolf, Technical Evangelist at Matrix, in this interview at @ThingsExpo, held Nov 4–6, 2014, at the Santa Clara Convention Center in Santa Clara, CA.
Nowadays, a large number of sensors and devices are connected to the network. Leading-edge IoT technologies integrate various types of sensor data to create a new value for several business decision scenarios. The transparent cloud is a model of a new IoT emergence service platform. Many service providers store and access various types of sensor data in order to create and find out new business values by integrating such data.
Too often with compelling new technologies market participants become overly enamored with that attractiveness of the technology and neglect underlying business drivers. This tendency, what some call the “newest shiny object syndrome,” is understandable given that virtually all of us are heavily engaged in technology. But it is also mistaken. Without concrete business cases driving its deployment, IoT, like many other technologies before it, will fade into obscurity.
There are so many tools and techniques for data analytics that even for a data scientist the choices, possible systems, and even the types of data can be daunting. In his session at @ThingsExpo, Chris Harrold, Global CTO for Big Data Solutions for EMC Corporation, will show how to perform a simple, but meaningful analysis of social sentiment data using freely available tools that take only minutes to download and install. Participants will get the download information, scripts, and complete end-to-end walkthrough of the analysis from start to finish. Participants will also be given the pract...
WebRTC services have already permeated corporate communications in the form of videoconferencing solutions. However, WebRTC has the potential of going beyond and catalyzing a new class of services providing more than calls with capabilities such as mass-scale real-time media broadcasting, enriched and augmented video, person-to-machine and machine-to-machine communications. In his session at @ThingsExpo, Luis Lopez, CEO of Kurento, will introduce the technologies required for implementing these ideas and some early experiments performed in the Kurento open source software community in areas ...
Electric power utilities face relentless pressure on their financial performance, and reducing distribution grid losses is one of the last untapped opportunities to meet their business goals. Combining IoT-enabled sensors and cloud-based data analytics, utilities now are able to find, quantify and reduce losses faster – and with a smaller IT footprint. Solutions exist using Internet-enabled sensors deployed temporarily at strategic locations within the distribution grid to measure actual line loads.