Shyam's Slide Share Presentations

VIRTUAL LIBRARY "KNOWLEDGE - KORRIDOR"

This article/post is from a third party website. The views expressed are that of the author. We at Capacity Building & Development may not necessarily subscribe to it completely. The relevance & applicability of the content is limited to certain geographic zones.It is not universal.

TO VIEW MORE CONTENT ON THIS SUBJECT AND OTHER TOPICS, Please visit KNOWLEDGE-KORRIDOR our Virtual Library

Showing posts with label .Shyamsunder Panchavati. Show all posts
Showing posts with label .Shyamsunder Panchavati. Show all posts

Tuesday, April 11, 2017

Emotional Intelligence Has 12 Elements. Which Do You Need to Work On? 04-12























Image credit : Shyam's Imagination Library



Esther is a well-liked manager of a small team. Kind and respectful, she is sensitive to the needs of others. She is a problem solver; she tends to see setbacks as opportunities. She’s always engaged and is a source of calm to her colleagues. Her manager feels lucky to have such an easy direct report to work with and often compliments Esther on her high levels of emotional intelligence, or EI. And Esther indeed counts EI as one of her strengths; she’s grateful for at least one thing she doesn’t have to work on as part of her leadership development. It’s strange, though — even with her positive outlook, Esther is starting to feel stuck in her career. She just hasn’t been able to demonstrate the kind of performance her company is looking for. So much for emotional intelligence, she’s starting to think.

The trap that has ensnared Esther and her manager is a common one: They are defining emotional intelligence much too narrowly. Because they’re focusing only on Esther’s sociability, sensitivity, and likability, they’re missing critical elements of emotional intelligence that could make her a stronger, more effective leader. A recent HBR article highlights the skills that a kind, positive manager like Esther might lack: the ability to deliver difficult feedback to employees, the courage to ruffle feathers and drive change, the creativity to think outside the box. But these gaps aren’t a result of Esther’s emotional intelligence; they’re simply evidence that her EI skills are uneven. In the model of EI and leadership excellence that we have developed over 30 years of studying the strengths of outstanding leaders, we’ve found that having a well-balanced array of specific EI capabilities actually prepares a leader for exactly these kinds of tough challenges.

There are many models of emotional intelligence, each with its own set of abilities; they are often lumped together as “EQ” in the popular vernacular. We prefer “EI,” which we define as comprising four domains: self-awareness, self-management, social awareness, and relationship management. Nested within each domain are twelve EI competencies, learned and learnable capabilities that allow outstanding performance at work or as a leader (see the image below). These include areas in which Esther is clearly strong: empathy, positive outlook, and self-control. But they also include crucial abilities such as achievement, influence, conflict management, teamwork and inspirational leadership. These skills require just as much engagement with emotions as the first set, and should be just as much a part of any aspiring leader’s development priorities.




For example, if Esther had strength in conflict management, she would be skilled in giving people unpleasant feedback. And if she were more inclined to influence, she would want to provide that difficult feedback as a way to lead her direct reports and help them grow. Say, for example, that Esther has a peer who is overbearing and abrasive. Rather than smoothing over every interaction, with a broader balance of EI skills she could bring up the issue to her colleague directly, drawing on emotional self-control to keep her own reactivity at bay while telling him what, specifically, does not work in his style. Bringing simmering issues to the surface goes to the core of conflict management. Esther could also draw on influence strategy to explain to her colleague that she wants to see him succeed, and that if he monitored how his style impacted those around him he would understand how a change would help everyone.

Similarly, if Esther had developed her inspirational leadership competence, she would be more successful at driving change. A leader with this strength can articulate a vision or mission that resonates emotionally with both themselves and those they lead, which is a key ingredient in marshaling the motivation essential for going in a new direction. Indeed, several studies have found a strong association between EI, driving change, and visionary leadership.

In order to excel, leaders need to develop a balance of strengths across the suite of EI competencies. When they do that, excellent business results follow.

How can you tell where your EI needs improvement — especially if you feel that it’s strong in some areas?

Simply reviewing the 12 competencies in your mind can give you a sense of where you might need some development. There are a number of formal models of EI, and many of them come with their own assessment tools. When choosing a tool to use, consider how well it predicts leadership outcomes. Some assess how you see yourself; these correlate highly with personality tests, which also tap into a person’s “self-schema.” Others, like that of Yale University president Peter Salovey and his colleagues, define EI as an ability; their test, the MSCEIT (a commercially available product), correlates more highly with IQ than any other EI test.

We recommend comprehensive 360-degree assessments, which collect both self-ratings and the views of others who know you well. This external feedback is particularly helpful for evaluating all areas of EI, including self-awareness (how would you know that you are not self-aware?). You can get a rough gauge of where your strengths and weaknesses lie by asking those who work with you to give you feedback. The more people you ask, the better a picture you get.

Formal 360-degree assessments, which incorporate systematic, anonymous observations of your behavior by people who work with you, have been found to not correlate well with IQ or personality, but they are the best predictors of a leader’s effectiveness, actual business performance, engagement, and job (and life) satisfaction. Into this category fall our own model and the Emotional and Social Competency Inventory, or ESCI 360, a commercially available assessment we developed with Korn Ferry Hay Group to gauge the 12 EI competencies, which rely on how others rate observable behaviors in evaluating a leader. The larger the gap between a leader’s self-ratings and how others see them, research finds, the fewer EI strengths the leader actually shows, and the poorer the business results.

These assessments are critical to a full evaluation of your EI, but even understanding that these 12 competencies are all a part of your emotional intelligence is an important first step in addressing areas where your EI is at its weakest. Coaching is the most effective method for improving in areas of EI deficit. Having expert support during your ups and downs as you practice operating in a new way is invaluable.

Even people with many apparent leadership strengths can stand to better understand those areas of EI where we have room to grow. Don’t shortchange your development as a leader by assuming that EI is all about being sweet and chipper, or that your EI is perfect if you are — or, even worse, assume that EI can’t help you excel in your career. 


View at the original source


Sunday, December 25, 2016

SRK to be conferred honorary doctorate by Hyderabad-based varsity 12-26






Hyderabad: Bollywood superstar Shah Rukh Khan is slated to be conferred with honorary doctorate by the city-based Maulana Azad National Urdu University tomorrow.  

President Pranab Mukherjee would be the chief guest at the sixth convocation of the university where nearly 48,000 students will be awarded degrees.

Telangana and Andhra Pradesh Governor ESL Narasimhan would be the guest of honour at the event to be presided over by the varsity’s Chancellor, Zafar Younus Sareshwala.

Telangana Deputy Chief Minister Mohammad Mahmood Ali will be the other guest of honour. 

Along with Shah Rukh Khan, Urdu aficionado and founder of Rekhta Foundation Rajiv Saraf would be given ‘honorius causa’ for their contribution in the promotion of Urdu language and culture, a release from the varsity said.

The university was established here in 1998.

About 2,885 graduates and postgraduates and 276 M.Phil and PhD scholars from various disciplines in regular courses will be awarded the degrees. Besides, 44,235 graduates and post graduates under the distance mode of learning would also be given degrees in absentia, the release added.

View at the original source


Wednesday, December 30, 2015

Children who sleep more get better grades 12-31


Children who sleep more get better grades.



Sleep plays a fundamental role in the way we learn. Emerging evidence makes a compelling case for the importance of sleep for language learning, memory, executive function, problem solving and behaviour during childhood.

A new study that my colleagues and I have worked on illustrated how an optimal quantity of sleep leads to more effective learning in terms of knowledge acquisition and memory consolidation. Poor quality of sleep – caused by lots of waking up during the night – has also been reported to be a strong predictor of lower academic performance, reduced capacity for attention, poor executive function and challenging behaviours during the day.

Many adolescents are sleep-deprived as they gain less sleep than the average recommended level – around nine hours for this group. But due to school commitments, teenagers are required to wake up early at a set time even if they have not achieved the optimal number of hours sleep.

Along with these early start times, teenagers also experience pubertal phase delay – meaning pubertal teenagers will sleep even less due to biological factors. Combined with late night activities, this can have a significant negative effect on the quality of sleep and therefore their behaviour during the day.
Insufficient and poor quality of sleep appear to be pervasive during adolescence. These can have various consequences such as an excessive daytime sleepiness, poor diet and in turn impairments in cognitive control, risk-taking behaviour, diminished control of attention and behaviour, as well as poor emotional control.

More sleep versus better sleep

In a recent study involving 48 students between 16 and 19-years-old recruited through an independent sixth form college in central London, my colleagues and I at the Lifespan Learning and Sleep Laboratory at UCL examined the link between sleep, academic performance and environmental factors.

Our results showed that the majority of the teenagers achieved just over seven hours of sleep, with an average bedtime at 11.37pm. Our study showed that a longer amount of sleep and earlier bedtimes – measures of sleep quantity – were most strongly correlated with better academic results obtained by the students on a number of tests taken at school. In contrast, measures that were indicative of sleep quality were mostly linked with students’ performances on verbal reasoning tests and on grade point averages on tests at school.

So it appears from our results that "longer sleep" is more closely related to academic performance, while "good night sleep" is more closely related to overall cognitive processing.

Why teens are getting less and less sleep

Our study also confirms findings from previous research showing that teenagers are getting at least two to three hours less sleep than is needed for their optimal brain development and a healthy lifestyle.

There are several modern lifestyle factors that have shown to impact on sleep. We found that consumption of energy drinks and coffee, and social media use half an hour before habitual bedtime were strongly associated with poorer sleep.

Our study has also shown that the negative impact of poor sleep on academic functioning is not always matched by a realisation of this fact by students themselves, therefore they may have little motivation to alter bad sleep habits. Unlike for adults, adolescence is a crucial time because of continual changes in the brain – so sleep is particularly important for a teenager’s health.

Conditions that can impact sleep

There is an added complexity to the sleep patterns of children with developmental disorders, despite the fact that they are more likely to suffer from sleep problems. So far, we have examined sleep, and cognitive and behavioural functioning in children with Down Syndrome, Williams Syndrome and ADHD. All our studies show that sleep has a very important impact on cognitive and daytime functioning of children with these conditions.

When we examined levels of sleep biomarkers – melatonin and cortisol – in children with Williams syndrome, a rare genetic disorder, it revealed that they had elevated levels of cortisol and dampened levels of melatonin. High cortisol and low melatonin levels before bedtime were strongly linked with delayed sleep onset – taking around 50 minutes in comparison to the typical 20 minutes to fall asleep.
Since cortisol is often described as a stress hormone, high levels of this hormone before bedtime may potentially cause sleep problems including difficulty in relaxing and falling asleep. This is an important result to consider before a child is prescribed a melatonin supplement – which might not be necessary to help solve their actual sleep problem.

The effects of the sleep disturbances extend beyond the individual. Parents of children with developmental disorders often experience heightened levels of stress and sleep problems because they are kept awake by their children.

All this shows how crucial it is for teenagers to get the right amount of sleep – otherwise it could have long-term impacts on their health and on their grades.

View at the original source

Monday, July 6, 2015

A Chip That Mimics Human Organs Is the Design of the Year 07-06

A Chip That Mimics Human Organs Is the Design of the Year




 Paola Antonelli, the senior curator of design and architecture at the Museum of Modern Art, added an intriguing object to the museum’s permanent collection. It was a clear plastic chip, no bigger than a thumb drive, and it could soon change the way scientists develop and test life-saving medicines.
Called Organs-On-Chips, it’s exactly what it sounds like: A microchip embedded with hollow microfluidic tubes that are lined with human cells, through which air, nutrients, blood and infection-causing bacteria could be pumped. These chips get manufactured the same way companies like Intel make the brains of a computer. But instead of moving electrons through silicon, these chips push minute quantities of chemicals past cells from lungs, intestines, livers, kidneys and hearts. 
Networks of almost unimaginably tiny tubes give the technology its name—microfluidics—and let the chips mimic the structure and function of complete organs, making them an excellent testbed for pharmaceuticals. The ultimate goal is to lessen dependence on animal test subjects and decrease time and cost for developing drugs. Last year, researchers from Harvard’s Wyss Institute for Biologically Inspired Engineering started a company called Emulate, which is now working with companies like Johnson & Johnson on just this idea: pre-clinical trial testing. The company is currently working on incorporating Emulate’s chips into its research and development programs.
When the Harvard team first published its findings on the chips in 2010, the research was purely scientific. Now, five years later, it’s not only been inducted into the world’s foremost design collection, it’s also been named Design of the Year.very year, London’s Design Museum names one project as the year’s best. Past winners have included Zaha Hadid’s ethically-questionable (but stunning) Heydar Aliyev Cultural Centre in Azerbaijan, a lightbulb, and a government website. That a piece of medical equipment developed by biological engineers is this year’s winner isn’t just a nod to the design’s worthiness. It also says something about how views of what counts as “design” are changing.
To be sure, Organs-On-Chips is aesthetically brilliant. Antonelli, who recently called synthetic biology the most exciting frontier in design, described the chips as the epitome of design innovation. “In some lucky cases, the form is striking,” she says, referring to objects born out of scientific research. “In this particular case, added bonus, not only is the form striking, but so is the function—the idea behind the object.” Like a biological system, the chip’s form dictates its function, and its form is undeniably beautiful. But that’s not the end of the story. “Most people say form follows function, but it’s exactly the opposite in biology,” says Donald Ingber, a bioengineer and founding director of the Wyss Institute, which developed the chip and is working on commercializing it. 

“Actually, that’s not fair. It’s a dynamic relationship.” The structure of a biological system will inevitably affect the way it works, but Ingber says the design principle works both ways. “If you change the function, you can actually modulate the structure,” he says, noting how the diameter of blood vessels will adapt to decrease the tension in people who develop hypertension.

Working on the microscale requires precision. The chip effectively replaces the three-dimensional structures of an organ—the renal tubules of a kidney, the alveoli of the lungs, the veins in a liver—with tissue-lined microfluidic channels. Then it emulates the mechanics of those structures. For example, running air through a channel while using a vacuum to introduce a flexing motion will simulate the patterns of human breathing. The chip’s translucent polymer, in which the channels are encased, allows scientists to see what’s happening inside organs on the microscale. The prototypes can also be linked together to form a whole-body network of organs.

Organs-On-Chips embraces the most basic of design principles: efficiency. “Design in its greatest simplicity is minimizing any system down to its elements so as to have the greatest impact,” says Ingber. Like an increasing number of researchers, he understands that good science requires an understanding of good design. The principles that govern the two fields aren’t totally separate—in fact, design is a thread that runs through every field. It’s heartening when a big award reminds us of that.

Saturday, May 2, 2015

This Guy Invented Shoes That Grow Five Sizes In Five Years For Kids In Developing Countries 05-02

 Shoes That Grow Five Sizes In Five Years For Kids In Developing Countries


“I had no idea how important shoes were,” founder Kenton Lee told BuzzFeed News.

Kenton Lee was working at an orphanage in Kenya when he noticed a little girl with the ends of her shoes cut off and her toes sticking out. It was then that he came up with the idea for The Shoe That Grows.


“For years the idea of these growing shoes wouldn’t leave my mind,” he told BuzzFeed News.
The first step was starting Because International with a few friends in 2006, a nonprofit devoted to “working with and helping those in extreme poverty,” their site says.
Kenton Lee
Proof of Concept
 

Lee and his team at first tried to give the idea to companies like Nike, Crocs, and Toms, to no avail. Eventually they found a “shoe development company” called Proof of Concept who agreed to help them with the design.


The shoe is made out of a high quality soft leather on top, and extremely durable rubber soles similar material to a tire, Lee said. They expand through a simple system of buckles, snaps, and pegs.
Proof of Concept
Because International
 

The shoes are predicted to last a minimum of five years, and expand five sizes in that time. The small size will fit preschoolers through fifth graders, while the large will fit fifth through ninth graders.

This Guy Invented Shoes That Grow Five Sizes In Five Years For Kids In Developing Countries
Via theshoethatgrows.org

“I had no idea how important shoes were before I went to Kenya,” Lee said. “But kids, especially in urban areas, can get infections from cuts and scrapes on their feet from going barefoot, and contract diseases that cause them to miss school.”


The 30-year-old, who started a church in Idaho with his wife, said he wanted to put these kids in the best possible position to succeed in their lives.
“If I can provide a kid with protection so they stay healthy and keep going to school, I’ll have done my part.”
Because International / Via becauseinternational.org

The shoes cost $10 a pair, and each pair goes into a “duffle bag” that can fit 50 pairs of shoes. Once one organization’s duffle bag is full, Because International ships it to the organization that flies with them to one of seven countries.


Donors can either buy shoes to distribute themselves, or buy a pair of shoes and choose one of five American nonprofit organizations to distribute them to orphanages and churches around the world.
Because International
Because International
 

So far about 2,500 children across seven countries are wearing the shoes, including in Ghana, Haiti, Peru, Colombia, and Kenya.


“We have about 500 left of our first order, currently being stored in a room in my house where my son sometimes chews on them,” said Lee, referring to his 11-month-old son (he also has another one the way). He said they have an order of 3,000 more pairs coming in July for people to donate.
The Shoe That Grows / Via theshoethatgrows.org

“We considered making even larger ones for teenagers,” Lee added, “but we were told that they didn’t want to wear ‘charity shoes,’ they wanted to wear something cooler.”


He said he’s now being flooded with requests, mostly from Americans, to make adult-sized versions.

Monday, April 27, 2015

How-to: Tune Your Apache Spark Jobs (Part 2) 04-28

How-to: Tune Your Apache Spark Jobs (Part 2)


In the conclusion to this series, learn how resource tuning, parallelism, and data representation affect Spark job performance.
In this post, we’ll finish what we started in “How to Tune Your Apache Spark Jobs (Part 1)”. I’ll try to cover pretty much everything you could care to know about making a Spark program run fast. In particular, you’ll learn about resource tuning, or configuring Spark to take advantage of everything the cluster has to offer. Then we’ll move to tuning parallelism, the most difficult as well as most important parameter in job performance. Finally, you’ll learn about representing the data itself, in the on-disk form which Spark will read (spoiler alert: use Apache Avro or Apache Parquet) as well as the in-memory format it takes as it’s cached or moves through the system.

Tuning Resource Allocation

The Spark user list is a litany of questions to the effect of “I have a 500-node cluster, but when I run my application, I see only two tasks executing at a time. HALP.” Given the number of parameters that control Spark’s resource utilization, these questions aren’t unfair, but in this section you’ll learn how to squeeze every last bit of juice out of your cluster. The recommendations and configurations here differ a little bit between Spark’s cluster managers (YARN, Mesos, and Spark Standalone), but we’re going to focus only on YARN, which Cloudera recommends to all users.
The two main resources that Spark (and YARN) think about are CPU and memory. Disk and network I/O, of course, play a part in Spark performance as well, but neither Spark nor YARN currently do anything to actively manage them.
Every Spark executor in an application has the same fixed number of cores and same fixed heap size. The number of cores can be specified with the --executor-cores flag when invoking spark-submit, spark-shell, and pyspark from the command line, or by setting the spark.executor.cores property in the spark-defaults.conf file or on aSparkConf object. Similarly, the heap size can be controlled with the --executor-cores flag or thespark.executor.memory property. The cores property controls the number of concurrent tasks an executor can run. --executor-cores 5 means that each executor can run a maximum of five tasks at the same time. The memory property impacts the amount of data Spark can cache, as well as the maximum sizes of the shuffle data structures used for grouping, aggregations, and joins.
The --num-executors command-line flag or spark.executor.instances configuration property control the number of executors requested. Starting in CDH 5.4/Spark 1.3, you will be able to avoid setting this property by turning ondynamic allocation with the spark.dynamicAllocation.enabled property. Dynamic allocation enables a Spark application to request executors when there is a backlog of pending tasks and free up executors when idle.
It’s also important to think about how the resources requested by Spark will fit into what YARN has available. The relevant YARN properties are:
  • yarn.nodemanager.resource.memory-mb controls the maximum sum of memory used by the containers on each node.
  • yarn.nodemanager.resource.cpu-vcores controls the maximum sum of cores used by the containers on each node.
Asking for five executor cores will result in a request to YARN for five virtual cores. The memory requested from YARN is a little more complex for a couple reasons: 
  • --executor-memory/spark.executor.memory controls the executor heap size, but JVMs can also use some memory off heap, for example for interned Strings and direct byte buffers. The value of thespark.yarn.executor.memoryOverhead property is added to the executor memory to determine the full memory request to YARN for each executor. It defaults to max(384, .07 * spark.executor.memory).
  • YARN may round the requested memory up a little. YARN’s yarn.scheduler.minimum-allocation-mb andyarn.scheduler.increment-allocation-mb properties control the minimum and increment request values respectively.
The following (not to scale with defaults) shows the hierarchy of memory properties in Spark and YARN:
And if that weren’t enough to think about, a few final concerns when sizing Spark executors:
  • The application master, which is a non-executor container with the special capability of requesting containers from YARN, takes up resources of its own that must be budgeted in. In yarn-client mode, it defaults to a 1024MB and one vcore. In yarn-cluster mode, the application master runs the driver, so it’s often useful to bolster its resources with the --driver-memory and --driver-cores properties.
  • Running executors with too much memory often results in excessive garbage collection delays. 64GB is a rough guess at a good upper limit for a single executor.
  • I’ve noticed that the HDFS client has trouble with tons of concurrent threads. A rough guess is that at most five tasks per executor can achieve full write throughput, so it’s good to keep the number of cores per executor below that number.
  • Running tiny executors (with a single core and just enough memory needed to run a single task, for example) throws away the benefits that come from running multiple tasks in a single JVM. For example, broadcast variables need to be replicated once on each executor, so many small executors will result in many more copies of the data.
To hopefully make all of this a little more concrete, here’s a worked example of configuring a Spark app to use as much of the cluster as possible: Imagine a cluster with six nodes running NodeManagers, each equipped with 16 cores and 64GB of memory. The NodeManager capacities, yarn.nodemanager.resource.memory-mb andyarn.nodemanager.resource.cpu-vcores, should probably be set to 63 * 1024 = 64512 (megabytes) and 15 respectively. We avoid allocating 100% of the resources to YARN containers because the node needs some resources to run the OS and Hadoop daemons. In this case, we leave a gigabyte and a core for these system processes. Cloudera Manager helps by accounting for these and configuring these YARN properties automatically.
The likely first impulse would be to use --num-executors 6 --executor-cores 15 --executor-memory 63G. However, this is the wrong approach because:
  • 63GB + the executor memory overhead won’t fit within the 63GB capacity of the NodeManagers.
  • The application master will take up a core on one of the nodes, meaning that there won’t be room for a 15-core executor on that node.
  • 15 cores per executor can lead to bad HDFS I/O throughput.
A better option would be to use --num-executors 17 --executor-cores 5 --executor-memory 19G. Why?
  • This config results in three executors on all nodes except for the one with the AM, which will have two executors.
  • --executor-memory was derived as (63/3 executors per node) = 21.  21 * 0.07 = 1.47.  21 – 1.47 ~ 19.

Tuning Parallelism

Spark, as you have likely figured out by this point, is a parallel processing engine. What is maybe less obvious is that Spark is not a “magic” parallel processing engine, and is limited in its ability to figure out the optimal amount of parallelism. Every Spark stage has a number of tasks, each of which processes data sequentially. In tuning Spark jobs, this number is probably the single most important parameter in determining performance.
How is this number determined? The way Spark groups RDDs into stages is described in the previous post. (As a quick reminder, transformations like repartition and reduceByKey induce stage boundaries.) The number of tasks in a stage is the same as the number of partitions in the last RDD in the stage. The number of partitions in an RDD is the same as the number of partitions in the RDD on which it depends, with a couple exceptions: thecoalescetransformation allows creating an RDD with fewer partitions than its parent RDD, the union transformation creates an RDD with the sum of its parents’ number of partitions, and cartesian creates an RDD with their product.
What about RDDs with no parents? RDDs produced by textFile or hadoopFile have their partitions determined by the underlying MapReduce InputFormat that’s used. Typically there will be a partition for each HDFS block being read. Partitions for RDDs produced by parallelize come from the parameter given by the user, orspark.default.parallelism if none is given.
To determine the number of partitions in an RDD, you can always call rdd.partitions().size().
The primary concern is that the number of tasks will be too small. If there are fewer tasks than slots available to run them in, the stage won’t be taking advantage of all the CPU available. 
A small number of tasks also mean that more memory pressure is placed on any aggregation operations that occur in each task. Any join, cogroup, or *ByKey operation involves holding objects in hashmaps or in-memory buffers to group or sort. join, cogroup, and groupByKey use these data structures in the tasks for the stages that are on the fetching side of the shuffles they trigger. reduceByKey and aggregateByKey use data structures in the tasks for the stages on both sides of the shuffles they trigger.
When the records destined for these aggregation operations do not easily fit in memory, some mayhem can ensue. First, holding many records in these data structures puts pressure on garbage collection, which can lead to pauses down the line. Second, when the records do not fit in memory, Spark will spill them to disk, which causes disk I/O and sorting. This overhead during large shuffles is probably the number one cause of job stalls I have seen at Cloudera customers.
So how do you increase the number of partitions? If the stage in question is reading from Hadoop, your options are:
  • Use the repartition transformation, which will trigger a shuffle.
  • Configure your InputFormat to create more splits.
  • Write the input data out to HDFS with a smaller block size.
If the stage is getting its input from another stage, the transformation that triggered the stage boundary will accept anumPartitions argument, such as
What should “X” be? The most straightforward way to tune the number of partitions is experimentation: Look at the number of partitions in the parent RDD and then keep multiplying that by 1.5 until performance stops improving. 
There is also a more principled way of calculating X, but it’s difficult to apply a priori because some of the quantities are difficult to calculate. I’m including it here not because it’s recommended for daily use, but because it helps with understanding what’s going on. The main goal is to run enough tasks so that the data destined for each task fits in the memory available to that task.
The memory available to each task is (spark.executor.memory * spark.shuffle.memoryFraction *spark.shuffle.safetyFraction)/spark.executor.cores. Memory fraction and safety fraction default to 0.2 and 0.8 respectively.
The in-memory size of the total shuffle data is harder to determine. The closest heuristic is to find the ratio between Shuffle Spill (Memory) metric and the Shuffle Spill (Disk) for a stage that ran. Then multiply the total shuffle write by this number. However, this can be somewhat compounded if the stage is doing a reduction:
Then round up a bit because too many partitions is usually better than too few partitions.
In fact, when in doubt, it’s almost always better to err on the side of a larger number of tasks (and thus partitions). This advice is in contrast to recommendations for MapReduce, which requires you to be more conservative with the number of tasks. The difference stems from the fact that MapReduce has a high startup overhead for tasks, while Spark does not.

Slimming Down Your Data Structures

Data flows through Spark in the form of records. A record has two representations: a deserialized Java object representation and a serialized binary representation. In general, Spark uses the deserialized representation for records in memory and the serialized representation for records stored on disk or being transferred over the network. There is work planned to store some in-memory shuffle data in serialized form.
The spark.serializer property controls the serializer that’s used to convert between these two representations. The Kryo serializer, org.apache.spark.serializer.KryoSerializer, is the preferred option. It is unfortunately not the default, because of some instabilities in Kryo during earlier versions of Spark and a desire not to break compatibility, but the Kryo serializer should always be used
The footprint of your records in these two representations has a massive impact on Spark performance. It’s worthwhile to review the data types that get passed around and look for places to trim some fat.
Bloated deserialized objects will result in Spark spilling data to disk more often and reduce the number of deserialized records Spark can cache (e.g. at the MEMORY storage level). The Spark tuning guide has a great section on slimming these down.
Bloated serialized objects will result in greater disk and network I/O, as well as reduce the number of serialized records Spark can cache (e.g. at the MEMORY_SER storage level.)  The main action item here is to make sure to register any custom classes you define and pass around using the SparkConf#registerKryoClasses API.

Data Formats

Whenever you have the power to make the decision about how data is stored on disk, use an extensible binary format like Avro, Parquet, Thrift, or Protobuf. Pick one of these formats and stick to it. To be clear, when one talks about using Avro, Thrift, or Protobuf on Hadoop, they mean that each record is a Avro/Thrift/Protobuf struct stored in asequence file. JSON is just not worth it. 
Every time you consider storing lots of data in JSON, think about the conflicts that will be started in the Middle East, the beautiful rivers that will be dammed in Canada, or the radioactive fallout from the nuclear plants that will be built in the American heartland to power the CPU cycles spent parsing your files over and over and over again. Also, try to learn people skills so that you can convince your peers and superiors to do this, too.