Showing posts with label SPSS. Show all posts
Showing posts with label SPSS. Show all posts

Wednesday, January 13, 2016

Getting Started with Data Science: Storytelling with Data

Earlier this month, IBM Press and Pearson have published my book titled: Getting Started with Data Science: Making Sense of Data with Analytics. You can download sample pages, including a complete chapter. There are 104 pages in the sample. You can also watch a brief interview about the book recorded earlier at the IBM Insight2015 Conference.

The very purpose of authoring this book was to rethink the way we have been teaching statistics and analytics to students and practitioners. It is no secret that most students required to take the mandatory stats course dislike it. I believe it has something to do with the way we have been teaching the subject than to do with the aptitude of our students. Furthermore, I believe there is a greater opportunity to equip the students with the skills needed in a world awash with data where competing on analytics defines the real competitive advantage.

No wonder, the latest issue of the leading publication on the subject, The American Statistician, is dedicated to reimagining how statistics should be taught in the undergraduate curriculum. The editors noted:
“We hope that this collection of articles as well as the online discussion provide useful fodder for further review, assessment, and continuous improvement of the undergraduate statistics curriculum that will allow the next generation to take a leadership role by making decisions using data in the increasingly complex world that they will inhabit.”
I am confident that my book will do its small part in equipping the next generation of students with the kind of skills needed to succeed in a data-centric world. For one, I have taken a storytelling approach to statistics. This book reinforces the point that data science and analytics training should be applied rather than theoretical, and the ultimate purpose of producing or consuming statistical analysis is to tell fascinating stories from it. Therefore, the book opens with the chapter titled, The Bazaar of Storytellers.

Who is this book for?

While the world is awash with large volumes of data, inexpensive computing power, and vast amounts of digital storage, the skilled workforce capable of analyzing data and interpreting it is in short supply. A 2011 McKinsey Global Institute report suggests that “the United States alone faces a shortage of 140,000 to 190,000 people with analytical expertise and 1.5 million managers and analysts with the skills to understand and make decisions based on the analysis of big data.”


Getting Started with Data Science (GSDS) is a purpose-written book targeted at those professionals who are tasked with analytics, but they do not have the comfort level needed to be proficient in data-driven analytics. GSDS appeals to those students who are frustrated with the impractical nature of the prescribed textbooks and are looking for an affordable text to serve as a long-term reference. GSDS embraces the 24-7 streaming of data and is structured for those users who have access to data and software of their choice, but do not know what methods to use, how to interpret the results, and most importantly how to communicate findings as reports and presentations in print or on-line.

GSDS is a resource for millions employed in knowledge-driven industries where workers are increasingly expected to facilitate smart decision-making using up-to-date information that sometimes takes the form of continuously updating data.

At the same time, the learning-by-doing approach in the book is equally suited for independent study by senior undergraduate and graduate students who are expected to conduct independent research for their coursework or dissertations.

Praise for the book

I am also pleased to share with you the praise for my book by Dr. Munir Sheikh, Canada’s former chief statistician:
“The power of data, evidence, and analytics in improving decision-making for individuals, businesses, and governments is well known and well documented. However, there is a huge gap in the availability of material for those who should use data, evidence, and analytics but do not know how. This fascinating book plugs this gap, and I highly recommend it to those who know this field and those who want to learn.”
— Munir A. Sheikh, Ph.D., Distinguished Fellow and Adjunct Professor at Queen’s University

Tom Davenport, author of the bestselling books Competing on Analytics and Big Data @ Work.has the following to say about my book:
“A coauthor and I once wrote that data scientists held ‘the sexiest job of the 21st century.’ This was not because of their inherent sex appeal, but because of their scarcity and value to organizations. This book may reduce the scarcity of data scientists, but it will certainly increase their value. It teaches many things, but most importantly it teaches how to tell a story with data.”
—Thomas H. Davenport, Distinguished Professor, Babson College; Research Fellow, MIT.

Dr. Patrick Surry
, Chief Data Scientist at www.Hopper.com had the following to say:
“This book addresses the key challenge facing data science today, that of bridging the gap between analytics and business value. Too many writers dive immediately into the details of specific statistical methods or technologies, without focusing on this bigger picture. In contrast, Haider identifies the central role of narrative in delivering real value from big data.

“The successful data scientist has the ability to translate between business goals and statistical approaches, identify appropriate deliverables, and communicate them in a compelling and comprehensible way that drives meaningful action. To paraphrase Tukey, ‘Far better an approximate answer to the right question, than an exact answer to a wrong one.’ Haider’s book never loses sight of this central tenet and uses many realworld examples to guide the reader through the broad range of skills, techniques, and tools needed to succeed in practical data-science. “Highly recommended to anyone looking to get started or broaden their skillset in this fast-growing field.”
And finally, Professor Atif Mian, author of the best-selling book: The House of Debt offered the following assessment:
“We have produced more data in the last two years than all of human history combined. Whether you are in business, government, academia, or journalism, the future belongs to those who can analyze these data intelligently. This book is a superb introduction to data analytics, a must-read for anyone contemplating how to integrate big data into their everyday decision making.”
— Professor Atif Mian, Theodore A. Wells ’29 Professor of Economics and Public Affairs,
Princeton University; Director of the Julis-Rabinowitz Center for Public Policy and Finance at the Woodrow Wilson School.

Wednesday, February 27, 2013

Workshops on Modelling Choices using R in Toronto

Making choices is inherently human. We choose between brands of cereal or amongst candidates in an election. At times, choices may be influenced by the characteristics of the decision maker, such as age, income and sex. Choices may also be influenced by the attributes of competing alternatives, such as the cost of travelling between two cities by air or rail. At other times, choices are influenced by both.

Analyzing choices can be tricky. Practitioners and researchers have developed numerous statistical techniques to analyze and model choices. This workshop will offer applied, hands-on training in analyzing choices.

The workshops will be offered in two sessions. First session will focus on binary (yes/no) choices and introduce the basic assumptions about choice analysis. It will provide hands-on training on exploratory data analysis. Second session will focus on advanced topics in choice modelling including multiple (multinomial) choices, elasticities, and estimating market shares.

Participants are expected to bring their own laptops. Basic concepts will be illustrated in SPSS, Stata, and R.

Title: Workshop on Modelling Choices

Dates and Time: Session One - Friday, March 22, 2013 (2pm-5pm)

Session Two - Friday, March 29, 2013 (2pm-5pm)

Instructor: Murtaza Haider, Ph.D.

Location: Ted Rogers School of Management, Ryerson University, 55 Dundas Street West, Room 3-119, Toronto M5G 2C3 

Registration fee: The workshop is sponsored by the Dean’s office at the Ted Rogers School of Management and is offered free-of-cost to the Ryerson community.

Please RSVP by emailing mba@ryerson.ca

Registration will be restricted to 25 participants.

Monday, July 30, 2012

Big data, big analytics, big opportunity

Data, data, every where
Nor any byte to think

The world today is awash with data. Corporations, governments, and individuals are busy generating petabytes of data on culture, economy, environment, religion, and society.  While data has become abundant and ubiquitous, data analysts needed to turn raw data into knowledge are in fact in short supply.

With big data comes big opportunity for the educated middle class in the developing world where an army of data scientists can be trained to support the offshoring of analytics from the western countries where such needs are unlikely to be filled from the locally available talent.

In a 2011 report, McKinsey Global Institute revealed that the United States alone faces a shortage of almost 200,000 data analysts. The American economy requires an additional 1.5 million managers proficient in decision-making based on insights gained from the analysis of large data sets. And even when Hal Varian, Google’s famed chief economist, profoundly proclaimed that “the real sexy job in 2010s is to be a statistician,” there were not many takers for the opportunity in the West where students pursuing degrees in statistics, engineering, and other empirical fields are small in number and are often visa students from abroad.

A recent report by Statistics Canada revealed that two-thirds of those who graduated with a PhD in engineering from a Canadian University in 2005 spoke neither English nor French as mother tongue. Similarly, four out of 10 PhD graduates in computers, mathematics, and physical sciences did not speak a western language as mother tongue. Also, more than 60 per cent of engineering graduates were visible minorities, suggesting that the supply chain of highly qualified professional talent in Canada, and to a large extent in North America, is already linked to the talent emigrating from China, Egypt, India, Iran, and Pakistan.

The abundance of data and the scarcity of analysts present a unique opportunity for developing countries, which have an abundant supply of highly numerate youth who could be trained and mobilized en masse to write a new chapter in offshoring. This would require a serious rethink for thought leaders in developing countries who have not taxed their imaginations beyond thinking of policies to create sweat shops where youth would undersell their skills and see their potential wilt away while creating undergarments for consumers in the west. The fate of the youth in developing countries need not be restricted to stitching underwear or making cold calls from offshored call-centers in order for them to be part of the global value chains. Instead, they can be trained as skilled number-crunchers who would add value to otherwise worthless data for businesses, big and small.

A multi-billion dollar industry

The past decade has witnessed a major change in the sectorial evolution of some very large manufacturing firms known in the past for mostly hardware engineering and now evolving into firms delivering services, such as business analytics. Take IBM for example, which specialized as a computer hardware company producing servers, desktop computers, laptops, and other supporting infrastructure. That was IBM’s past. Today, IBM is focused on analytics. It is spending hundreds of millions of dollars in advertising, trying to rebrand itself as a leader in business analytics. In fact, it has divested from several hardware initiatives, such as manufacturing laptops, and has instead spent billions in acquisitions to build its analytic credentials. For instance, IBM has acquired SPSS for over a billion dollars to capture the retail side of the Business analytics market. For large commercial ventures, IBM acquired Cognos to offer full service analytics.

In 2011 alone, the business analytics software market was worth over $30 billion. Oracle ($6.1bn), SAP ($4.6 bn), IBM ($4.4 bn), and Microsoft and SAS each with $3.3 bn in sales led the market. It is estimated that the sale of business analytics software alone will hit $50 billion by 2016.  Dan Vesset of IDC, a company specializing in watching industry trends, aptly noted that business analytics had “crossed the chasm into the mainstream mass market” and the “demand for business analytics solutions is exposing the previously minor issue of the shortage of highly skilled IT and analytics staff.”

In addition to the bundled software and service sales offered by the likes of Oracle and IBM, business analytics services in the consulting domain generated several billion dollars more worldwide. While the large firms command the lion’s share in the analytics market, the billions left at the bottom are still a large enough prize to take the analytics plunge.

Several billion reasons to hop on the analytics bandwagon

While the IBMs of the world are focused largely on large corporations, the analytics needs for small and medium-sized enterprises (SMEs) are unlikely to be met by IBM, Oracle, or other large players. Cost is the most important determinant. SMEs prefer to have analytics done on the cheap while the overheads of the large analytics firms run into millions of dollars thus pricing them out of the SME market. With offshoring comes the access to affordable talent in developing countries who can bid for smaller contracts and beat the competition in the West on price, and over time on quality as well.

The trick therefore, is to beat the IBMs of the world in the analytics game by not competing against them. Realizing that business analytics is not a market, but an amalgamation of several types of markets focused on delivering value-added services involving data capture, data warehousing, data cleaning, data mining, and data analysis, developing countries can carve out a niche for themselves by focusing exclusively on contracts that large firms will not bid for because of their intrinsic large overheads.

Leaving the fight for top dollars in analytics to top dogs, a cottage industry in analytics could be developed in the developing countries that may strive to serve the analytics need of SMEs. Take the example of the Toronto Transit Commission (TTC), Canada’s largest public transit agency with annual revenues exceeding a billion dollars. When TTC needed to have a large database of almost a half million commuter complaints analyzed, it turned to Ryerson University, rather than a large analytics firm. TTC’s decision to work with Ryerson University was motivated by two considerations. First the cost; as a public sector university, Ryerson believes strongly in serving the community and thus offered the services for gratis. The second reason is quality. Ryerson University, like most similar institutions of higher learning, excels in analytics where several faculty members work at the cutting edge of analytics and are more than willing to apply their skills to real life problems.

Why now?

The timing had never been better to undertake such an endeavor on a very large scale. The innovations in Information and Communication Technology (ICT) and the ready availability of the most advanced analytics software as freeware allows entrepreneurs in developing countries to compete worldwide. The Internet makes it possible to be part of global marketplaces with negligible costs. With cyber marketplaces such as Kijiji and Craigslist individuals can become proprietors offering services worldwide.

Using the freely available Google Sites, one can have a business website online immediately at no cost.Google Docs, another free service from Google, allows one to have a web server for free to share documents with collaborators or the rest of the world for free. Other free services, such as Google Trends, allow individual researchers to generate data on business and social trends without needing subscriptions for services that cost millions. The graph below is generated using Google trends showing daily visits to the websites of leading analytics firms. Without free access to such services, access to the data used to generate the same graph would carry a huge price tag.

Similarly, another free service from Google allows one to determine, for instance, which cities registered the highest number of search requests for ‘business analytics’. It appears that four of the top six cities where analytics are most popular are located in India, which is evident from the following graph where search intensity is mapped on a normalized index of 0 to 100.

The other big development of recent times is freeware that is leveling the playing field between haves and have-nots. In analytics, one of the most sophisticated computing platforms is R, which is available for free. Developers worldwide are busy developing the R platform, which now offers over 3,000 packages for free for analyzing data. From econometrics to operations research, R is fast becoming the lingua franca for computing. R has evolved from being popular just amongst computing geeks to having its praise sung by the New York Times.

R has also made some new friends, especially Paul Butler, a Canadian student who became a worldwide sensation by mapping the geography of Facebook. While being an intern at Facebook, Paul analyzed gigabytes of data to plot how Facebook’s friends were linked globally. His map (see the image below) became an instant hit worldwide and has been reproduced in publications thousands of times. If you are wondering what software Paul used to generate the map, wonder no more, the answer is R.

R is fast becoming the preferred computing platform for data scientists worldwide. For decades the data analysis market was ruled by the likes of SAS, SPSS, Stata and other similar players. R has taken over the imagination of data analysts as of late who are fast converging to R, especially after R’s ability to interact with Hadoop (another open source platform) for analyzing big data . In fact, most innovations in statistics are first coded in R so that the algorithms become available to all immediately and for free.

Source: http://r4stats.com/articles/popularity/

The fact that R is freely available should not be taken lightly. A commercial license of a similarly equipped version of SPSS may cost up to US$7,500. The other big advantage of using R is the fact that thousands of training documents on the Internet and videos on YouTube are also available for free by volunteers.

Where to next

The private sector has to take the lead for business analytics to take root in developing countries. The governments could also have a small role in regulation. However, the analytics revolution has to take place not because of the public sector, but in spite of it. Even public sector universities in developing countries cannot be entrusted with the task where senior university administers do not warm up to innovative ideas unless they involve a junket in Europe or North America. At the same time the faculty in public sector universities in developing countries is often unwilling to try new technologies.

The private sector in developing countries may want to launch first an industry group that takes upon the task of certifying firms and individuals interested in analytics for quality, reliability, and ethical and professional competencies. This will help build confidence around national brands. Without such certification, foreign clients will be apprehensive to share their proprietary data with individuals hidden behind computer monitors thousands of miles away.

The private sector will also have to take the lead in training a professional workforce in analytics. Several companies train their employees in the latest technology and then market their skills to clients. The training houses would therefore also double as consulting practices where the best graduates may be retained as consultants.

Small virtual marketplaces could be setup in large cities where clients can put requests for proposals and pre-screened, qualified bidders can compete for the contract. The national self-regulating body will be responsible for screening qualified bidders from its vendor-of-record database, which it would make available to clients globally through the Internet.

The IBMs of the world see the analytics market to hit hundreds of billions in revenue in the next decade. The abundant talent in developing countries can be polished into a skilled workforce to tap into the analytics market to channel some revenue to developing countries while creating gainful employment opportunities for the educated youth who have been reduced to making cold calls from offshored call centers.

Saturday, February 12, 2011

Q: A new software for analyzing survey data

Q is a new market research software from Australia that specializes in the analysis of market research surveys. Q has been designed for analysts who primarily work with survey data. I test drove the professional edition of Q and found it to be a welcome addition.

I downloaded Q from http://www.q-researchsoftware.com/download.aspx and obtained a free 30-day trial license, which offers full functionality, including importing one’s own data for analysis. I imported a few data sets into Q without any hassle. Q reads SPSS as well as CSV files.

Q has a tabbed GUI with four distinct tabs. The opening tab is called Table, which by default presents the summary statistics of the first numeric variable in the data set. The second tab is called Variables and Questions, which in fact is the most important tab because here variables are designated a particular type, i.e., categorical, continuous, etc., which then determines what analytics could be performed using a particular
variable.

clip_image002

Since Q is specifically designed for analysing survey data, variables are also designated by the type of answer solicited in the actual survey instrument. Thus variables are categorized as ‘pick one’ in instances where respondents were presented with multiples choices from which they had to pick one choice. Similarly, other types include ‘pick any’, ‘number grid’, or ‘ranking’. The type ‘experiment’ refers to variables that capture conjoint analysis data in a stated preference choice experiment.

The Data tab presents a tabular view of the underlying data and the Notes tab allows one to review notes related with the data set. The Q GUI also informs the analyst if the data were filtered or weighted using a weight variable. The filtering option allows simple as well as complex filtering that may involve one or more variables.

The first variable in my data set was the ID variable that identified each respondent in the data set. Q by default computed descriptive statistics for the ID variable. Since I was not particularly interested in determining the average value for the ID variable, I selected a categorical variable from my data, and Q quickly displayed the frequency table as percentages. When I changed the Summary option and instead opted for another categorical variable, Q quickly displayed a crosstab between the two variables.

For stated preference data, Q uses the wide data format where each respondent is represented in a single row. Most econometrics software use the long format for stated preference data because it prevents one from managing a large number of variables. Consider working with a stated preference data about the choice of airlines where the survey respondents are presented with five alternatives (i.e., airlines). Let us further assume that each alternative is identified by three attributes, such as airfare, flight duration, and on-flight amenities. Let us also assume that each attribute, such as price, has three levels: low, medium, and high. And lastly, let us assume that each respondent is presented with five sets of choices, which Q refers to as tasks.

This experiment will generate 225 variables (5 x 3 x 3 x 5), which are sometimes hard to manage, though Q offers a sophisticated environment to setup the experiment as described above in the Variables and Questions tab. Also, when the raw stated preference data are recovered from computerized survey instruments that automatically populate the database, such as Sawtooth or web-based survey tools, the data are already in wide format, which Q can handle with ease.

Strengths

Q offers certain unique features that are not available in other software. When it presents crosstabs, it adds arrows to suggest statistical significance such that the blue arrow represents positive significance and a red arrow suggests negative significance. In the choice experiments, the software presents estimated coefficients from the model in tabular format representing different coefficients for explanatory variables. A crosstab between airline brand type and trip purpose is presented in Figure 2. The significance test reported by Q is not testing the null hypothesis that the estimated coefficient equals 0. Instead, Q compares the significance of the estimated coefficient for one trip purpose against its compliment. Notice the last row which shows a blue coloured upward arrow for the price coefficient for business travellers against the red coloured downward arrow for the price coefficient for holiday travellers. The coloured coefficients suggest that they are statistically different from each other, where the business traveller appear to be less price-sensitive than the holiday travellers.

clip_image004

While most econometrics software allow one to test if the estimated coefficients for different groups are statistically different from each other, the process requires several extra steps. Q on the other hand does it on the fly.

Another unique feature is called Banner, which allows complex crosstabs involving more than two categorical variables. For instance, consider one wants to determine if the preference for a particular brand differs by gender, age groups, and country of residence. Q would present these differences in a single table whereas most other statistical analysis software would generate multiple crosstabs. Furthermore, Q permits one to aggregate categories by clicking on the output in a crosstab, thus eliminating the need to first recode the variable.

Q is also well integrated with Excel. I was able to export tables from Q into Excel with a simple mouse click. Q automatically formatted the same table in Excel and generated a graph from the same table a separate sheet. This further simplifies sharing results with colleagues who may not work with Q and therefore can review the results in MS Excel. Q also comes with a free version that allows one to review Q’s output.

clip_image006

Weaknesses

Q is distinct from other software in many ways. However, some features in Q are very unique and do not conform to the intuitive base knowledge, which most analyst have usually accumulated by working with other software in the past, such as SPSS and Excel. One key distinction is Q’s unique nomenclature. Q calls a simple regression model with a categorical explanatory variable ‘split cell experiment’. The main disadvantage of its unique nomenclature is that most new users of Q, who may have worked with other similar software or have taken courses in statistics/ market research, would have no exposure to Q’s unique terminology, which therefore has to be learnt afresh.

Q supports a point and click environment and does not generate a log or a syntax file. This makes reproducing results or repeating the analysis a more cumbersome task. Perhaps the developers may want to include this feature in a later version.

Q is rather expensive for the analytics it offers. Q professional costs $1,499 per license. A transferable license costs three times as much. Within advanced analytics Q supports OLS, Generalized least squares, logit models, cluster analysis, and principal component analysis. There are whole host of other advanced econometric tools, which are available in other competing software that cost much less.

Final word

Whereas Q has many unique features, most of its advanced core competencies are readily available in other software, such as SPSS, and Stata, Eviews, and R. It does stand out in offering advanced data managing capabilities for survey data. Q will be a preferred tool for those market researchers who rely more on cross tabulations. For others who subject their data to advanced econometrics, such as nested logit models, testing for self-selection biases, or post-estimation tests, Q offers a rather restricted set of tools.
Enhanced by Zemanta