Skip to main content
Self-Service Analytics offers connection to Cloudera’s open source Hadoop platform - Cloudera Distributed Hadoop (CDH)*. CDH provides unified batch processing, interactive SQL, interactive search, and role-based access controls. In addition, it offers enterprise-grade continuous availability. Specifically, Self-Service Analytics connects to CDH’s fault‐tolerant storage system called the Hadoop Distributed File System (HDFS). The Self-Service Analytics HDFS connector uses its own embedded Apache Spark functionality. It supports Apache Spark 2.2 in its implementation. By default, the HDFS connector is not included with Self-Service Analytics. You or your administrator need to download and enable it before configuring the connector. After setting up the connector, create data sources that specify the necessary connection information and identify the data you want to use. See Create and Manage Data Sources for more information. After you set up your data sources, create dashboards and visuals from the data in these data sources.

Feature Support

Connector support for specific features is shown in the following table. Key: Y - Supported; N - Not Supported; N/A - not applicable