Diploma

Big Data and Data Science Diploma

Computer Science & AI
Location
Online
Duration
1 Year

The Big Data and Data Science (BD&DS) Professional Diploma equips participants with the skills to collect, manage, and analyze large-scale, complex datasets using modern data science methods and tools. Through four specialized courses and an applied project, participants gain hands-on experience turning raw data into actionable insights, preparing them to meet the growing demands of research and industry in an increasingly data-driven world.

Starts

Starts

Oct, 2026

Credential

Credential

NU Certificate

Location

Location

Online

Language

Language

English
(Arabic Support)

Program overview

The Big Data and Data Science (BD&DS) Professional Diploma is a structured professional program designed to provide participants with a comprehensive foundation in the principles, methodologies, and applications of Big Data and Data Science. The program comprises four specialized courses and an applied project, through which participants are expected to integrate and demonstrate the knowledge, analytical competencies, and practical skills acquired throughout their studies. The courses are carefully selected from the Informatics program offerings within the School of Information Technology and computer science to ensure academic rigor, professional relevance, and alignment with current developments and practices in the field.

What You Will Learn

Understand the architecture of Hadoop clusters at both the hardware and system software levels.

Apply Hadoop and related Big Data technologies such as MapReduce, Spark, Hive, Impala, and Pig in developing analytics and solving the types of problems faced by enterprises today.

Use advanced technological tools to analyze data.

Understand key concepts in data mining and Big Data data analytics.

Solve real-life problems using state-of-the-art technologies developed for cloud-based machine learning computing and data analyses.

Curriculum

Big Data and Data Science Diploma Courses

The capability of collecting and storing huge amounts of versatile data necessitate the development and use of new techniques and methodologies for processing and analyzing big data. The Big Data landscape is continuously evolving as new technologies emerge and existing technologies mature. This is a comprehensive course covering Spark and key elements of the Hadoop Ecosystem used in developing end to end applications for processing Big Data efficiently. Students who complete this course will understand key Spark and Hadoop concepts, and they will learn to apply Spark and Hadoop tools in developing applications for solving the types of problems faced by enterprises and research institutions today.

This course provides an introduction to machine learning and statistical data analysis.  The course provides an introduction to the basic probability theory, statistics, and statistical data analysis. Topics such as parameter estimation, hypothesis testing and regression analysis will be covered in the course. In addition, the course will focus on machine learning topics including: Bayes classifiers, K-nn, decision trees, SVM, K-means, principal component analysis, independent component analysis and Neural Nets.

During this course, students will learn how to solve real-life problems using state-of-the-art technologies developed for cloud-based machine learning computing and data analyses.

This is an applied course where students can develop on their combined knowledge of Big Data technologies (e.g. Hadoop, Spark, etc.) and Data Science (e.g. Statistics, Machine Learning, etc.) and understand how such combination is used to solve real-world applications. In addition to this main goal, the course has the additional goal of familiarizing students with the latest technological and scientific trends in the field and how Big Data and data science are used in modern business enterprises. Use cases of real problems such as networking traffic, text analytics, and financial applications will be addressed in this course.

This introductory course focuses on building the requisite understanding and skills necessary to start applying Deep Learning models. Special emphasis will be convolutional architectures, recurrent neural network and autoencoders. We will attempt varied applications areas, from image classification, to text generation, to signal processing.

Data Sciences is a fast evolving practice that apply several sciences, theories and techniques in solving complex data-related problems and develop applications that support transforming the way organizations do their activities. Machine learning techniques are at the core of data sciences and have proven great value in the context of the practical applications of Big Data solutions. In this course, the students will be able to link the machine learning theories and methods in a practical real-life use case context. The hands-on labs will reinforce the concepts learned from the Introduction to Machine Learning course (CIT-651) with deeper focus in applying them to enable customers realize business value.

Furthermore, students will learn how the actual customer engagements in this filed works including consulting and implementation approaches. The course will demonstrate tools and technologies – both open source and commercial – like R, Java/Spring, Weka, Hadoop, Spark, Giraph, Cloud Foundry, Madlib and Greenplum applied to practical situations

After such long learning journey of evolutionally growing Big Data technologies and Data Science techniques and applications; the objective of the group project is to put all what students have learned during the 4 courses of the diploma into a real-life end to end customer-like engagement to strengthen the expertise they have gained and acquired through their contribution over the two semesters of the diploma. 

Under mentoring provided by the project supervisors, each group of typically 5 students will select a project topic and start applying industry driven CRISP-DM lifecycle to build end to end data driven use case. Over 6 weeks of mentorship, each group will follow key milestones to produce final solution and present their work for discussion and evaluation.

Admission Requirements

Solid understanding of enterprise information systems and applications.

Working experience with programming and applications development preferably with java and python.

Understating of probability, statistics, linear algebra, and math competence.

Solid understanding of data warehousing, business intelligence, data analysis and its applications.

Comfortable with python for data structures and analysis including commonly used packages - background in other tools.

37,500 EGP
Application Fees (Non Refundable)
1,500 EGP
Price Per Course
7,500 EGP

Have Questions?