Interview with Dr. Gregory Piatetsky-Shapiro, Data Science Pioneer and KDnuggets President

I recently had the great privilege of interviewing one of the most renowned experts in data science and machine learning: Dr. Gregory Piatetsky-Shapiro. Gregory is the president of KDnuggets, a leading site on AI, analytics, big data, data mining, and data science, with over 200,000 monthly visitors. He is also the co-founder of ACM SIGKDD, the premier professional organization for knowledge discovery and data mining.

Gregory has been a pioneer in the field since its early days. After studying computer science at NYU and earning his PhD in 1984 with a thesis on applying machine learning to databases, he went on to become a researcher, software developer, and chief scientist at several startups.

But his biggest impact has come from his volunteer work building the data science community. Gregory created and organized the first Knowledge Discovery in Databases (KDD) workshops from 1989 to 1994, which evolved into the annual KDD conference. He also launched the KDnuggets email newsletter in 1993, before the days of the web, and it has since grown into a hugely popular resource site.

Over the course of his distinguished career, Gregory has worn many hats – researcher, practitioner, consultant, organizer, and communicator. Speaking with him provided an enlightening look at how far data science has come and where it may be heading.

The Early Days of Data Mining and KDD

Gregory was involved in machine learning going back to his days as a graduate student at NYU in the early 1980s. For his PhD thesis, he worked on applying ML to databases and information retrieval.

"This was a time before the term ‘data mining‘ or ‘data science‘ really existed," he recalled. "The focus was more on knowledge discovery and machine learning, but a lot of the fundamental concepts were taking shape."

In 1989, Gregory organized a workshop on Knowledge Discovery in Databases (KDD) at the IJCAI conference in Detroit. This was the first workshop devoted to KDD and it brought together a diverse group of researchers from AI, databases, statistics, and other fields.

Gregory went on to organize the first three KDD workshops in 1989, 1991, and 1993. Attendance and interest grew each time. After the 1993 workshop, he decided to recruit Usama Fayyad to organize the next event in 1994.

Together with Ramasamy Uthurusamy, they transformed the 1995 edition into a full-fledged conference sponsored by AAAI. However, they realized that KDD was an interdisciplinary endeavor and could benefit from a broader base. In 1998, they worked with database researcher Won Kim to establish KDD as an official Special Interest Group (SIG) within the ACM.

The annual KDD conference remains the premier research event in the field to this day. Attendance has grown dramatically over the years:

Year # Attendees
1995 400
2000 500
2005 700
2010 900
2015 1800
2020 3300 (virtual)

(Sources: Gregory Piatetsky-Shapiro, ACM SIGKDD, KDD Conferences)

Meanwhile, the KDnuggets newsletter that Gregory launched in 1993 to connect the fledgling KDD community has evolved into a leading industry website. From just 50 subscribers in the beginning, it has grown to over 200,000 unique monthly visitors and boasts popular social media accounts on Twitter and LinkedIn.

The AI and Deep Learning Revolution

A major theme of my conversation with Gregory was the remarkable progress in artificial intelligence and machine learning capabilities in recent years, driven in large part by the rise of deep learning.

He pointed to benchmarks showing machines surpassing human-level performance on tasks like voice recognition, image classification, and language translation. For example, Microsoft achieved a 5.9% word error rate in conversational speech recognition in 2017, beating the 5.1% human parity threshold. Google brought its English-Chinese translation BLEU score up to 44.8 that same year, matching average human translators.

"There has been amazing progress in just the last few years," Gregory said. "Capabilities that many people thought were impossible are now becoming reality. It‘s an exciting time but also one that raises profound questions for society."

While he doesn‘t lose sleep over far-fetched doomsday scenarios of AI turning against humanity, Gregory does worry about the disruptive impact of automation on the job market and economic equality.

He envisions AI being able to handle increasingly sophisticated cognitive tasks, potentially displacing knowledge workers at a large scale. Autonomous vehicles alone could eliminate millions of driving jobs. An oft-cited 2013 study by Oxford researchers estimated that 47% of US jobs were at high risk of computerization.

At the same time, Gregory is optimistic about the beneficial potential of AI to solve major challenges and improve quality of life. He cited promising applications in areas like healthcare (personalized treatment, drug discovery), education (adaptive learning, intelligent tutoring), transportation (reduced congestion and emissions), and scientific research (analysis of massive datasets).

"My hope is that we can harness the power of AI as a tool to expand human capacity, not replace it," he said. "But this will require proactive social policies to support displaced workers and ensure the gains are shared broadly. I think data science itself has a big role to play in designing fair, effective solutions."

The Evolving Data Science Toolbox

Another perennial hot topic in data science is the choice of programming languages and tools. I asked Gregory to weigh in on the Python vs R debate and reflect on the evolution of the field.

Based on KDnuggets polls, both Python and R have seen steady growth in adoption as primary languages for analytics and data science. In the latest 2022 survey, Python was used by 61% of respondents and R by 26%, while traditional tools like SAS, SPSS, and MATLAB have declined to under 10% combined.

"Python has emerged as the most popular and versatile platform for data science and ML," Gregory observed. "It‘s easier to learn than R and has an extensive ecosystem of libraries like pandas, scikit-learn, and TensorFlow. But R still has advantages for statistics and visualization. And tools like RStudio are adding support for Python."

He noted that Python‘s rise has been fueled by the growth of deep learning, as the major frameworks have prioritized Python APIs. It has also benefited from increased adoption in academia and integration into data engineering pipelines.

Looking ahead, Gregory expects Python‘s lead to continue growing but sees a role for a diversity of tools:

"Data science is a team sport, and different tools fit different niches. Python is great as a Swiss army knife, but R, SQL, Java, and even Excel remain important. I think the next frontier is in making ML more accessible with low-code/no-code platforms."

Advice for Aspiring Data Scientists

For students and professionals looking to enter the burgeoning field of data science, Gregory offered some insights based on his experience and observations.

First, he emphasized the importance of building strong foundations in mathematics (especially statistics, probability, and linear algebra), computer science fundamentals, and domain expertise. Programming skills are essential, but data science is not just about coding.

"I see a lot of aspiring data scientists dive straight into the hottest algorithms without understanding the underlying principles," he cautioned. "But having that solid grasp of math and stats is what separates true data scientists from mere technicians."

At the same time, he stressed the value of hands-on experience working with data and noted that there are increasingly accessible points of entry.

"You don‘t need a PhD to get started in data science," he said. "There are so many great online courses, boot camps, and competitions now. Find a project you‘re passionate about, whether it‘s entering a Kaggle contest, analyzing some open data, or automating a process at work. The best way to learn is by doing."

When it comes to landing a data science job, Gregory highlighted the importance of showcasing your skills through a portfolio of projects, open source contributions, and a compelling online presence.

"Your GitHub and personal website speak louder than your resume these days," he advised. "Hiring managers want to see your code, your thought process, how you communicate insights."

He also emphasized the value of engaging with the data science community through events, online forums, and platforms like KDnuggets, Analytics Vidhya, and local meetup groups.

"Networking and continuous learning are so crucial in this fast-moving field," he said. "Surround yourself with brilliant, passionate people and soak up as much knowledge as you can. Never stop learning."

The Bright Future of Data Science in India

As an Indian data scientist myself, I had to ask Gregory for his perspective on the field‘s trajectory in my country.

He noted that India already has an impressive bench of data science talent, second only to the United States in visitors to KDnuggets. Indian universities and technology firms have long excelled in mathematics, statistics, and computer science.

"India‘s strengths in quantitative fields and programming are a natural fit for data science," he said. "Combined with the entrepreneurial spirit and growing digital economy, I see immense potential for India to be a leading hub of data-driven innovation."

He pointed to the vibrant data science communities on Analytics Vidhya and other platforms as a grassroots force shaping the Indian ecosystem. But he also acknowledged the need for broader data literacy and translation of data science from academia to industry.

"There is still a gap between the top-tier talent and the mainstream," he said. "Democratizing data skills and tying them to real business needs will be key to India realizing its full potential in this space."

Overall though, he expects India‘s role in the global data science community to continue rising in the coming years. He even quipped that KDnuggets may need to expand its editorial team in Bengaluru!

Cultivating Data Science Community

As a final topic, I wanted to get Gregory‘s reflections on the role of community and knowledge-sharing in data science. Both KDnuggets and Analytics Vidhya have made this a core part of our missions.

"Data science is inherently collaborative and thrives on openness," he said. "Unlike more mature fields, a lot of the cutting-edge work is happening out in the open on platforms like GitHub, Kaggle, and arXiv. We‘re all learning and building on each other‘s work."

He pointed to the volunteer efforts behind the KDD conferences, KDnuggets, and other initiatives as examples of the ethos of paying it forward. Despite being a prominent figure, he still sees himself as a student of the craft.

"I‘ve been fortunate to play a part in growing the community, but I‘ve gained far more than I‘ve given," he said humbly. "The contributors and readers are the real stars."

Moving forward, Gregory hopes to see data science become even more accessible and inclusive. He dreams of a world where data literacy is widespread and people from all backgrounds can participate in harnessing data for good.

"We‘re not there yet, but I‘m optimistic," he said. "As the philosopher Eric Hoffer wrote, ‘In times of change, learners inherit the earth, while the learned find themselves beautifully equipped to deal with a world that no longer exists.‘ In data science, we are all perpetual learners."

I couldn‘t agree more, and I left our conversation feeling grateful and energized to continue advancing that mission. I hope some of Gregory‘s wisdom will rub off on me and on all of you in the Analytics Vidhya community!

This article is based on an interview conducted with Gregory Piatetsky-Shapiro in March 2023.

How useful was this post?

Click on a star to rate it!

Average rating 0 / 5. Vote count: 0

No votes so far! Be the first to rate this post.

Similar Posts