Shyam's Slide Share Presentations

VIRTUAL LIBRARY "KNOWLEDGE - KORRIDOR"

This article/post is from a third party website. The views expressed are that of the author. We at Capacity Building & Development may not necessarily subscribe to it completely. The relevance & applicability of the content is limited to certain geographic zones.It is not universal.

TO VIEW MORE CONTENT ON THIS SUBJECT AND OTHER TOPICS, Please visit KNOWLEDGE-KORRIDOR our Virtual Library

Showing posts with label Data Science. Show all posts
Showing posts with label Data Science. Show all posts

Tuesday, May 30, 2017

Data science: teaming skills to harness insights faster 05-30



What skills determine success or failure for the new leaders of this age: data science teams? 



Data is the currency of the new millennium. But distilling and simplifying that data to gain real insight requires an increasingly complex skill set and equally sophisticated tools. Gartner researcher Peter Sondergaard sums up the power of data analysis in the context of other notable innovations in history stating, “Information is the oil of the 21st century, and analytics is the combustion engine.”
If you break down the origins of data, you’ll find that 20 percent of the world’s data is public, while the other 80 percent is proprietary. But, like any powerful tool, it needs a lead or guiding light, and data scientists have quickly risen to this challenge with developers and data engineers as their collaborators. These groups join virtually and physically to learn, curate, build and deploy analytic solutions to help extract insights from their vast data stores. This new cross functional collaboration has helped the data unit function as one, elevating the role of the data science team within the larger enterprise.

As organizations become increasingly data-driven and the influence of the data scientist skyrockets, what are the soft skills that determine success or failure for the new A-Team?



  • Creativity and imagination: The best data science teams are patient, persistent and focused. They understand how data pipelines function and are able to identify alternate solutions if something goes awry. The members of a data science team also love to learn, and their curiosity helps them come up with unexpected fixes to problems. When the different roles in a data science team come together, the result is a combined knowledge of numerous types of data sets and different programming languages.

  • Rigor and discipline: Data science teams manage enormous amounts of data every day. A good understanding of procedures and standards is crucial to stay on top of it all. When each member of the data science team is clear on best practices, the data management process is streamlined, therefore making life easier for those who rely on data to do their jobs. With a firm grasp on algorithms, code and how it benefits the infrastructure, data science teams have exponential power within their organizations.

  • Business acumen: Data science teams are the foundation of any data-driven organization. As such, they need to have a holistic view into how the business operates and what problems the company is looking to solve. A successful data science team has this information at their fingertips so they know how the data will ultimately be used to propel the organization towards the larger organization’s goals.
Once the data science team is built with these traits and proper guidelines are in place, a technological infrastructure with flexibility at its core must be created. Today’s organizations store data on a combination of public cloud, private cloud and on-premise hardware. Data science teams must be able to consistently manage data no matter where it is stored. In addition, because every industry has its own unique processes and compliance standards that data science tools must incorporate, the platforms themselves should be easily customizable.

Consider an actual example from IBM, NASA and the SETI Institute. These organizations are working together to analyze more than six terabytes of complex deep space radio signals to hunt for patterns that might identify the presence of intelligent extraterrestrial life. With the proper tools—IBM Analytics on Apache Spark, part of the Data Science Experience—SETI has been able to embark on its Stellar Pair Eavesdropping campaign, which enables the organization to look for potential communications between planets that might be orbiting in double star systems. More than half of all stars are, in fact, these types of planets. By extracting new features from millions of observations, researchers are able to use machine-learning techniques to classify signals and sharpen their focus for subsequent deep analysis on clusters of signals which are anomalous or outliers.

Without high-performing data science professionals and the right collaboration tools, organizations like SETI would not be able to handle and ultimately realize the full potential of their data. Just as an artist requires different tools for different creations, a data scientist needs a palette of capabilities to resolve the different problems they need to solve. IBM’s data science environment offers the most advanced analytics, open source technology and integrated development community, all built to encourage creativity and collaboration.


View at the original source


Wednesday, April 20, 2016

Just How Smart Are Smart Machines? 04-20


Just How Smart Are Smart Machines? 


The number of sophisticated cognitive technologies that might be capable of cutting into the need for human labor is expanding rapidly. But linking these offerings to an organization’s business needs requires a deep understanding of their capabilities.






If popular culture is an accurate gauge of what’s on the public’s mind, it seems everyone has suddenly awakened to the threat of smart machines. Several recent films have featured robots with scary abilities to outthink and manipulate humans. In the economics literature, too, there has been a surge of concern about the potential for soaring unemployment as software becomes increasingly capable of decision making. Yet managers we talk to don’t expect to see machines displacing knowledge workers anytime soon — they expect computing technology to augment rather than replace the work of humans. In the face of a sprawling and fast-evolving set of opportunities, their challenge is figuring out what forms the augmentation should take. Given the kinds of work managers oversee, what cognitive technologies should they be applying now, monitoring closely, or helping to build?

To help, we have developed a simple framework that plots cognitive technologies along two dimensions. (See “What Today’s Cognitive Technologies Can — and Can’t — Do.”) First, it recognizes that these tools differ according to how autonomously they can apply their intelligence. On the low end, they simply respond to human queries and instructions; at the (still theoretical) high end, they formulate their own objectives. Second, it reflects the type of tasks smart machines are being used to perform, moving from conventional numerical analysis to performance of digital and physical tasks in the real world. The breadth of inputs and data types in real-world tasks makes them more complex for machines to accomplish.

By putting those two dimensions together, we create a matrix into which we can place all of the multitudinous technologies known as “smart machines.” More important, this helps to clarify today’s limits to machine intelligence and the challenges technology innovators are working to overcome next. Depending on the type of task a manager is targeting for redesigned performance, this framework reveals the various extents to which it might be performed autonomously and by what kinds of machines.


Four Levels of Intelligence


Clearly, the level of intelligence of smart machines is increasing. The general trend is toward greater autonomy in decision making — from machines that require a highly structured data and decision context to those capable of deciphering a more complex context.


Support for Humans


For decades, the prevailing assumption has been that cognitive technologies would provide insight to human decision makers — what used to be known as “decision support.” Even with IBM Corp.’s Watson and many of today’s other cognitive systems, most people assume that the machine will offer a recommended decision or course of action but that a human will make the final decision.


Repetitive Task Automation


It is a relatively small step to go from having machines support humans to having the machines make decisions, particularly in structured contexts. Automated decision making has been gaining ground in recent years in several domains, such as insurance underwriting and financial trading; it typically relies on a fixed set of rules or algorithms, so performance doesn’t improve without human intervention. Typically, people monitor system performance and fine-tune the algorithms.


Context Awareness and Learning


Sophisticated cognitive technologies today have some degree of real-time contextual awareness. As data flow more continuously and voluminously, we need technologies that can help us make sense of the data in real time — detecting anomalies, noticing patterns, and anticipating what will happen next. Relevant information might include location, time, and/or a user’s identity, which might be used to make recommendations (for example, the best route to work based on the time of day, current traffic levels, and the driver’s preference for highways versus back roads).

One of the hallmarks of today’s cognitive computing is its ability to learn and improve performance. Much of the learning takes place through continuous analysis of real-time data, user feedback, and new content from text-based articles. In settings where results are measurable, learning-oriented systems will ultimately deliver benefits in the form of better stock trading decisions, more accurate driving time predictions, and more precise medical diagnoses.


Self-Awareness


So far, machines with self-awareness and the ability to form independent objectives reside only in the realm of fiction. With substantial self-awareness, computers may eventually gain the ability to work beyond human levels of intelligence across multiple contexts, but even the most optimistic experts say that general intelligence in machines is three to four decades away.


Four Cognitive Task Types


A straightforward way to sort out tasks performed by machines is according to whether they process only numbers, text, or images — the building blocks of cognition — or whether they know enough to take informed actions in the digital or physical world.


Analyzing Numbers


The root of all cognitive technologies is computing machines’ superior performance at analyzing numbers in structured formats (typically, rows and columns). Classically, this numerical analysis was applied purely in support of human decision makers. People continued to perform the front-end cognitive tasks of creating hypotheses and framing problems, as well as the back-end interpretation of the numbers’ implications for decisions. Even as analysts added more visual analytics displays and more predictive analytics in the past decade, people still did the interpretation.

Today, companies are increasingly embedding analytics into operational systems and processes to make repetitive automated decisions, which enables dramatic increases in both speed and scale. And whereas it used to take a human analyst to develop embedded models, “machine learning” methods can produce models in an automated or semiautomated fashion.


Analyzing Words and Images


A key aspect of human cognition is the ability to read words and images and to determine their meaning and significance. But today, a wide variety of technological tools, such as machine learning, natural language processing, neural networks, and deep learning, can classify, interpret, and generate words. Some of them can also analyze and identify images.

The earliest intelligent applications involving words and images involved text, image, and speech recognition to allow humans to communicate with computers. Today, of course, smartphones “understand” human speech and text and can recognize images. These capabilities are hardly perfect, but they are widely used in many applications.

When words and images are analyzed on a large scale, this comprises a different category of capability. One such application involves translating large volumes of text across languages. Another is to answer questions as a human would. A third is to make sense of language in a way that can either summarize it or generate new passages.

IBM Watson was the first tool capable of ingesting, analyzing, and “understanding” text well enough to respond to detailed questions. However, it doesn’t deal with structured numerical data, nor can it understand relationships between variables or make predictions. It’s also not well suited for applying rules or analyzing options on decision trees. However, IBM is rapidly adding new capabilities included in our matrix, including image analysis.

There are other examples of word and image systems. Most were developed for particular applications and are slowly being modified to handle other types of cognitive situations. Digital Reasoning Systems Inc., for example, a company based in Franklin, Tennessee, that developed cognitive computing software for national security purposes, has begun to market intelligent software that analyzes employee communications in financial institutions to determine the likelihood of fraud.

Another company, IPsoft Inc., based in New York City, processes spoken words with an intelligent customer agent programmed to interpret what customers want and, when possible, do it for them.
IPsoft, Digital Reasoning, and the original Watson all use similar components, including the ability to classify parts of speech, to identify key entities and facts in text, to show the relationships among entities and facts in a graphical diagram, and to relate entities and relationships with objectives. This category of application is best suited for situations with much more — and more rapidly changing — codified textual information than any human could possibly absorb and retain.

Image identification and classification are hardly new. “Machine vision” based on geometric pattern matching technology has been used for decades to locate parts in production lines and read bar codes. Today, many companies want to perform more sensitive vision tasks such as facial recognition, classification of photos on the Internet, or assessment of auto collision damage. Such tasks are based on machine learning and neural network analysis that can match particular patterns of pixels to recognizable images.

The most capable machine learning systems have the ability to “learn” — their decisions get better with more data, and they “remember” previously ingested information. For example, as Watson is introduced to new information, its reservoir of information expands. Other systems in this category get better at their cognitive task by having more data for training purposes. But as Mike Rhodin, senior vice president of business development for IBM Watson, noted, “Watson doesn’t have the ability to think on its own,” and neither does any other intelligent system thus far created.

Performing Digital Tasks

One of the more pragmatic roles for cognitive technology in recent years has been to automate administrative tasks and decisions. In order to make automation possible, two technical capabilities are necessary. First, you need to be able to express the decision logic in terms of “business rules.” Second, you need technologies that can move a case or task through the series of steps required to complete it. Over the past couple of decades, automated decision-making tools have been used to support a wide variety of administrative tasks, from insurance policy approvals to information technology operations to high-speed trading.

Lately, companies have begun using “robotic process automation,” which uses work flow and business rules technology to interface with multiple information systems as if it were a human user. Robotic process technology has become popular in banking (for back-office customer service tasks, such as replacing a lost ATM card), insurance (for processing claims and payments), information technology (IT) (for monitoring system error messages and fixing simple problems), and supply chain management (for processing invoices and responding to routine requests from customers and suppliers).

The benefits of process automation can add up quickly. An April 2015 case study at Telefónica O2, the second-largest mobile carrier in the United Kingdom, found that the company had automated over 160 process areas using software “robots.” The overall three-year return on investment was between 650% and 800%.


Performing Physical Tasks


Physical task automation is, of course, the realm of robots. Though people love to call every form of automation technology a robot, one of Merriam-Webster’s definitions of robot is “a machine that can do the work of a person and that works automatically or is controlled by a computer.”

In 2014, companies installed about 225,000 industrial robots globally, more than one-third of them in the automotive industry. However, robots often fall well short of expectations. In 2011, the founder of Foxconn Technology Co., Ltd., a Taiwan-based multinational electronics contract manufacturing company, said he would install one million robots within three years, replacing one million workers. However, the company found that employing only robots to build smartphones was easier said than done. To assemble new iPhone models in 2015, Foxconn planned to hire more than 100,000 new workers and install about 10,000 new robots.

Historically, robots that replaced humans required a high level of programming to do repetitive tasks. For safety reasons, they had to be segregated from human workers. However, a new type of robots — often called “collaborative robots” — can work safely alongside humans. They can be programmed simply by having a human move their arms.

Robots have varying degrees of autonomy. Some, such as remotely piloted drone aircraft and robotic surgical instruments and mining equipment, are designed to be manipulated by humans. Others become at least semiautonomous once programmed but have limited ability to respond to unexpected conditions. As robots get more intelligence, better machine vision, and increased ability to make decisions, they will integrate other types of cognitive technologies while also having the ability to transform the physical environment. IBM Watson software, for example, has been installed in several different types of robots.


The Great Convergence


Slowly but surely, the worlds of artificially intelligent software and robots seem to be converging, and the boundaries between different cognitive technologies are blurring. In the future, robots will be able to learn and sense context, robotic process automation and other digital task tools will improve, and smart software will be able to analyze more intricate combinations of numbers, text, and images.
We anticipate that companies will develop cognitive solutions using the building blocks of application program interfaces (APIs). One API might handle language processing, another numerical machine learning, and a third question-and-answer dialogue. While these elements would interact with each other, determining which APIs are required will demand a sophisticated understanding of cognitive solution architectures.

This modular approach is the direction in which key vendors are moving. IBM, for example, has disaggregated Watson into a set of services — a “cognitive platform,” if you will — available by subscription in the cloud. Watson’s original question-and- answer services have been expanded to include more than 30 other types, including “personality insights” to gauge human behavior, “visual recognition” for image identification, and so forth. Other vendors of cognitive technologies, such as Cognitive Scale Inc., based in Austin, Texas, are also integrating multiple cognitive capabilities into a “cognitive cloud.”

Despite the growing capabilities of cognitive technologies, most organizations that are exploring them are starting with small projects to explore the technology in a specific domain. But others have much bigger ambitions. For example, Memorial Sloan Kettering Cancer Center, in New York City, and the University of Texas MD Anderson Cancer Center, in Houston, Texas, are taking a “moon shot” approach, marshaling cognitive tools like Watson to develop better diagnostic and treatment approaches for cancer.

 Continued on page 2


Designing a Cognitive Architecture



























Sunday, April 19, 2015

Data science demands elastic infrastructure 04-19

Data science demands elastic infrastructure




Those companies that try to run big data projects in data centers may be setting themselves up for failure. Matt Asay explains. 

As companies struggle to make sense of their increasingly big data, they're laboring to figure out the morass of technologies necessary to become successful. However, many will remain stymied, because they keep trying to fit a necessarily fluid process of asking questions of one's data with outmoded, rigid data infrastructure.
Or as Amazon Web Services (AWS) data science chief Matt Wood tells it, they need the cloud.
While the cloud isn't a panacea, its elasticity may well prove to be the essential ingredient to big data success.

How much cloud do I need?

The problem with trying to run big data projects within a data center revolves around rigidity. As Matt Wood told me in a recent interview, this problem "is not so much about absolute scale of data but rather relative scale of data."
In other words, as a company's data volume takes a step function up or down, enterprise infrastructure can't keep up. In his words, "Customers will tool for the scale they're currently experiencing," which is great... until it's not.
In a separate conversation, he elaborates:
"Those that go out and buy expensive infrastructure find that the problem scope and domain shift really quickly. By the time they get around to answering the original question, the business has moved on. You need an environment that is flexible and allows you to quickly respond to changing big data requirements. Your resource mix is continually evolving--if you buy infrastructure, it's almost immediately irrelevant to your business because it's frozen in time. It's solving a problem you may not have or care about any more."
Success in big data depends upon iteration, upon experimentation as you try to figure out the right questions to ask and the best way to answer them. This is hard when dealing with a calcified infrastructure.

A eulogy for the data center?

Of course, it's not quite so simple as "all cloud, all the time."
Data, it would seem, has to obey fundamental laws of gravity, as Basho CTO Dave McCrory told TechRepublic in an interview:
"Big data workloads will live in large data centers where they are most advantaged. Why will they live in specific places? Because data attracts data.
"If I already have a large quantity of data in a specific cloud, I'm going to be inclined to store additional quantities of large data in the same place. As I do this and add workloads that interact with this data, more data will be created."
Over time, enterprises will look to the public cloud for all the reasons Wood describes, but legacy data is unlikely to make the migration. There's simply no reason to try to house old data in new infrastructure. Not most of the time.
But some companies will find that they're more comfortable with existing data centers and will eschew the cloud. I'm not talking about hide-bound enterprise curmudgeons that shout "Phooey!" every time AWS is mentioned, either. No, sometimes the most data center-centric of companies will be the innovators like Etsy.
As Etsy CTO Kellan Elliott-McCrea informed TechRepublic, once Etsy had "gained confidence" in its ability to manage its Hadoop clusters (and other technology), they brought them in-house, netting a 10X increase in utilization and "very real cost savings."
Nor is Etsy alone. Other new-school web companies like Twitter have opted to run their own data centers, finding that this gives them greater control over their data.

You're no Twitter

As highly as you may estimate your abilities, the reality is that you're probably not an Etsy, Twitter, or Google. As painful as it is to say it, most of us are average. By definition.
This is what Microsoft's great genius was: rather than cater to the Übermensch of IT, Microsoft lowered the bar to becoming productive as a system administrator, developer, etc. In the process, Microsoft banked billions in profits, helping make a good sysadmin better or a decent developer good.
Regardless, all enterprises need to establish infrastructure that helps them to iterate. Some, like Etsy, may have figured out how to do this in their data centers--but for most of us, most of the time, Wood's advice rings true: "You need an environment that is flexible and allows you to quickly respond to changing big data requirements."
In other words, odds are that you're going to need the cloud.

Wednesday, February 11, 2015

Big Data and Data Center Operations 02-12


Big Data and Data Center Operations



Scott Koegler recently wrote a post titled ”Use Big Network Data to Predict and Avoid Network Problems” where he describes the use of data analysis and predictive analytics by IPSoft. Scott wrote:
"By turning predictive analytics inward to track where breaks happen most frequently, IT and network admins can set more accurate thresholds for recurring issues, positioning themselves ahead of any damage.”
Big data and predictive analytics are a great fit for the data center and IT operations, especially within the modern data center. There’s a great deal of data being generated and managed within IT operations, and the approaches and systems found within big data can help better understand and manage operations.
The example that Scott (and IPSoft) used is one that can provide value for any organization because it allows IT operations to understand (and predict) when breaks or issues might arise within the data center, network or remote office.
Imagine how impressive it would be for an IT specialist in a central office to get a notification that something in a far-flung branch office is amiss. Imagine again how that notification could tell IT staff exactly what was wrong and provide a recommendation to “fix” the problem before it actually became a problem.  This capability is available today with predictive analytics and data analysis.
Using predictive analytics and other big data approaches to identify bottlenecks, manage incidents and fix issues faster is the next logical step for data center and IT operations.
View at the original source

Tuesday, December 23, 2014

‘Data Scientist’ Replaces ‘Social Media Scientist’ In LinkedIn’s 2014 Top Skills List 12-23

‘Data Scientist’ Replaces ‘Social Media Scientist’ In LinkedIn’s 2014 Top Skills List - 




What a difference one short year makes. In 2013, my position of ‘Social Media Scientist’ which focuses on ‘social media marketing’ for clients and publishers was the most viable job in the land. This year, not only did ‘Data Scientist’ take the top spot in Linkedin’s “Top 25 Job Skills” list, the social media marketing skill was wiped clean off the 2014 list entirely.

Social Media Marketing no longer needed?


Does that mean my position has suddenly become irrelevant? Not really! What it says to me is that data collection, mining analysis and Cloud management are now the number-one recruitment darlings.

While us social networking folks were sought after last year, in 2014, those that slice and dice data came up more often in the 330 million member profiles. They, plus those involved in middleware, integration software, storage systems management and information security are the jobs in greater demand today.

Data Tech Gets its Sea Legs


This makes sense based on how many companies decided this past year to store their data in the Cloud -- whether it be the Internet of Things, smart homes, wearable devices, or TV streaming. All of these professional fields just matriculated from geek to chic - from hobby to mainstream vocation - so much so that TV tracking experts at Nielsen decided to add online television viewing to its area of purview.

Netflix & Amazon Prime Get Big Boy Pants


Nielsen Media, the company which has conducted standard TV measurement ever since that technology first graced our living rooms back in the 1950’s - has just announced plans to measure subscription online video services such as Netflix and Amazon Prime.

While both services have long been mum on ratings for both acquired programs and original series, Nielsen now plans to tabulate the data for them using that which was collected for clients and studios to see how their acquired content is performing on both Netflix and Amazon.

Data Scientists now Sexy


Just a quick perusal of articles posted by the likes of FortuneHarvard, Gigaom and Time Magazine gave more spotlight time to Data Scientists as well. 
‘Data Stars’ like Andrea Burbank of Pinterest, Patrick Poels of EventBrite, Silvanus Lee of Dropbox and Surabhi Gupta of Airbnb made the 
Fortune list, while former Netflix Senior Data Scientist Mohammad Sabah noted his position at the streaming video platform required the capture and analysis of an incredible amount of data  -- trying to figure out what consumers’ 
preferences were and what they wanted to view next. 
At the time, Sabah said, 75 percent of users selected movies based on the company’s recommendations, and based on those findings, Netflix has sought to make that number escalate even higher.

To that end, Netflix’s most-interesting use of data might be its attempts to actually analyze what’s going on in movies themselves. Sabah said it had already captured JPEGs and noted the exact time that credits start rolling, and it also took into account other characteristics. For instance, It could make a lot of sense out of variables, such as volume, colors and scenery that might give valuable signals about what viewers really like.

What this means for finding jobs in 2015?


"People call them (data scientists) unicorns" because the combination of skills required is so rare, said Jonathan Goldman, who ran LinkedIn Corp.'s team that in 2007 developed the "People You May Know" button -- which five years later drove more than half of the invitations on the professional-networking platform.

Employers according to Goldman says the ideal candidate must have more than traditional market-research skills: the ability to find patterns in millions of pieces of data streaming in from different sources, to infer from those patterns how customers behave and to write statistical models that pinpoint behavioral triggers.

Anyone with "data science" in his or her job title on a LinkedIn page is going to get "100 recruiter emails a day," said Josh Sullivan, who leads a 500-person data-science group at the consulting firm Booz Allen Hamilton Holding Corp.

So Says the Social Media Scientist. . .


So with this somewhat disturbing news that my field of play is not on the top of mind awareness for recruitment, I am content in the fact, I’ve secured a solid client base over the years.

On the other hand, ironically data shows that social media still does matter. 

Ninety-two percent of business owners today recognize the need to invest in social media as a necessary part of their business strategy, as it relates to engaging and marketing to their consumers; that’s an increase from 2013, when that number was 86%.

And as newer social networking platforms surface with ‘payment models' such as Tsu, Bitlanders, MyLifeB, BonzoMe and Bubblews - in my estimation that’s a whole new crop of networks that are going to need Social Media Scientists like myself to manage, market and promote for SMBs as well as larger brands. For more on this topic, see my 
previous post titled, “2015 Prediction For Top Digital Job: ‘Paid Social Specialist.’”  

That’s enough work to keep me busy for at least another year — and who knows when Linkedin’s “Top 25 Job Skills” list rolls around in 2015, my skill set might just be back in high demand, once again!

Happy job hunting and recruiting in 2015 -- you, Digital and 
Social Media Scientists!