Showing 31 total results
Learn to process, analyze, and derive insights from massive datasets using advanced tools and technologies.
The course covers essential concepts like lifecycle phases, repositories, and artifact management, ensuring efficient and reproducible builds.
It covers essential concepts such as RDDs, DataFrames, and SparkSQL for efficient data manipulation and querying.
It covers core concepts such as stream processing, stateful computations, event time processing, and fault tolerance.
It covers essential topics such as data ingestion, transformation, routing, and delivery across diverse systems.
It covers key components such as HDFS (Hadoop Distributed File System), YARN (Yet Another Resource Negotiator), and MapReduce for parallel data processing.
Participants will gain hands-on experience in leveraging Spark for real-world data analytics and machine learning applications.
It covers key concepts such as Kafka architecture, producers, consumers, topics, partitions, and brokers.
It covers core concepts such as Resilient Distributed Datasets (RDDs), DataFrames, and SparkSQL, providing the foundation for data manipulation and analysis.
This course provides system administrators with the skills necessary to set up, configure, manage, and troubleshoot Apache Spark clusters.
It covers essential topics like connecting to data sources, designing charts, and building custom dashboards for data exploration.
It covers various technologies, including Hadoop, Apache Spark, and NoSQL databases.
The training covers the concepts of hubs, links, and satellites to design a robust data architecture.
The course covers topics such as load balancing, reverse proxy setup, and configuring server blocks to handle multiple websites.
The course covers essential topics like setting up virtual hosts, managing server modules, and securing the server environment.
Participants will explore foundational principles, key components, and best practices for creating scalable, efficient, and secure data architectures.
The training focuses on building proficiency in using Talend Studio for data integration, transformation, and managing data workflows.
This course is suitable for data engineers, data scientists, and cloud professionals aiming to leverage Spark's powerful distributed computing capabilities in cloud environments.
Participants will learn the core Kafka architecture, installation and configuration, cluster administration, topic management, monitoring, and best practices for reliability and security.
Apache Superset basics and with deeper understanding of advanced configuration, custom visualizations, data security, performance optimization, and embedding Superset in enterprise applications.
This training provides a comprehensive understanding of Apache ZooKeeper, a centralized service for maintaining configuration information, naming, synchronization, and group services in distributed systems.
Apache Camel, a powerful open-source integration framework that enables seamless communication between different systems using Enterprise Integration Patterns (EIPs).
This course provides administrators with in-depth knowledge of Apache Superset, an open-source data exploration and visualization platform.
Participants will learn how to leverage data analytics to improve governance, enhance citizen services, and enable evidence-based decision-making.
Participants will learn how to install, configure, and administer Kylin, build and optimize data models, and perform interactive analytics on large datasets using SQL on Hadoop.
Participants will learn how to set up, configure, manage, and optimize ActiveMQ for reliable asynchronous communication in distributed systems.
The course focuses on writing and optimizing MapReduce programs, understanding Hadoop’s data flow, and integrating with other ecosystem tools such as Hive, Pig, and HBase
It focuses on core administration skills, including deployment, configuration, monitoring, and troubleshooting.
This training program is designed to help developers and testers leverage Apache Maven for managing and executing automated tests.
This training provides a practical and project-based introduction to Spark NLP, a production-grade natural language processing library built on Apache Spark.
This course provides hands-on training in developing and deploying modular OSGi applications using Apache Karaf.
This course introduces Apache Iceberg, an open table format for managing large analytic datasets on data lakes.