Databricks Databricks-Certified-Data-Analyst-Associate - Databricks Certified Data Analyst Associate Exam

Databricks Databricks-Certified-Data-Analyst-Associate Premium Access Download Demo

Page: 1 / 2
Total 65 questions

A data analyst has created a Query in Databricks SQL, and now they want to create two data visualizations from that Query and add both of those data visualizations to the same Databricks SQL Dashboard.

Which of the following steps will they need to take when creating and adding both data visualizations to the Databricks SQL Dashboard?

They will need to alter the Query to return two separate sets of results.

They will need to add two separate visualizations to the dashboard based on the same Query.

They will need to create two separate dashboards.

They will need to decide on a single data visualization to add to the dashboard.

They will need to copy the Query and create one data visualization per query.

Question # 2

A data analyst has created a user-defined function using the following line of code:

CREATE FUNCTION price(spend DOUBLE, units DOUBLE)

RETURNS DOUBLE

RETURN spend / units;

Which of the following code blocks can be used to apply this function to the customer_spend and customer_units columns of the table customer_summary to create column customer_price?

SELECT PRICE customer_spend, customer_units AS customer_price FROM customer_summary

SELECT price FROM customer_summary

SELECT function(price(customer_spend, customer_units)) AS customer_price FROM customer_summary

SELECT double(price(customer_spend, customer_units)) AS customer_price FROM customer_summary

SELECT price(customer_spend, customer_units) AS customer_price FROM customer_summary

Question # 3

A stakeholder has provided a data analyst with a lookup dataset in the form of a 50-row CSV file. The data analyst needs to upload this dataset for use as a table in Databricks SQL.

Which approach should the data analyst use to quickly upload the file into a table for use in Databricks SOL?

Create a table by uploading the file using the Create page within Databricks SQL

Create a table via a connection between Databricks and the desktop facilitated by Partner Connect.

Create a table by uploading the file to cloud storage and then importing the data to Databricks.

Create a table by manually copying and pasting the data values into cloud storage and then importing the data to Databricks.

Question # 4

Consider the following two statements:

Statement 1:

Statement 2:

Which of the following describes how the result sets will differ for each statement when they are run in Databricks SQL?

The first statement will return all data from the customers table and matching data from the orders table. The second statement will return all data from the orders table and matching data from the customers table. Any missing data will be filled in with NULL.

When the first statement is run, only rows from the customers table that have at least one match with the orders table on customer_id will be returned. When the second statement is run, only those rows in the customers table that do not have at least one match with the orders table on customer_id will be returned.

There is no difference between the result sets for both statements.

Both statements will fail because Databricks SQL does not support those join types.

When the first statement is run, all rows from the customers table will be returned and only the customer_id from the orders table will be returned. When the second statement is run, only those rows in the customers table that do not have at least one match with the orders table on customer_id will be returned.

Explanation:

Based on the images you sent, the two statements are SQL queries for different types of joins between the customers and orders tables. A join is a way of combining the rows from two table references based on some criteria. The join type determines how the rows are matched and what kind of result set is returned. The first statement is a query for a LEFT SEMI JOIN, which returns only the rows from the left table reference (customers) that have a match with the right table reference (orders) on the join condition (customer_id). The second statement is a query for a LEFT ANTI JOIN, which returns only the rows from the left table reference (customers) that have no match with the right table reference (orders) on the join condition (customer_id). Therefore, the result sets for the two statements will differ in the following way:

The first statement will return a subset of the customers table that contains only the customers who have placed at least one order. The number of rows returned will be less than or equal to the number of rows in the customers table, depending on how many customers have orders. The number of columns returned will be the same as the number of columns in the customers table, as the LEFT SEMI JOIN does not include any columns from the orders table.

The second statement will return a subset of the customers table that contains only the customers who have not placed any order. The number of rows returned will be less than or equal to the number of rows in the customers table, depending on how many customers have no orders. The number of columns returned will be the same as the number of columns in the customers table, as the LEFT ANTI JOIN does not include any columns from the orders table.

The other options are not correct because:

A. The first statement will not return all data from the customers table, as it will exclude the customers who have no orders. The second statement will not return all data from the orders table, as it will exclude the orders that have a matching customer. Neither statement will fill in any missing data with NULL, as they do not return any columns from the other table.

C. There is a difference between the result sets for both statements, as explained above. The LEFT SEMI JOIN and the LEFT ANTI JOIN are not equivalent operations and will produce different outputs.

D. Both statements will not fail, as Databricks SQL does support those join types. Databricks SQL supports various join types, including INNER, LEFT OUTER, RIGHT OUTER, FULL OUTER, LEFT SEMI, LEFT ANTI, and CROSS. You can also use NATURAL, USING, or LATERAL keywords to specify different join criteria.

E. The first statement will not return only the customer_id from the orders table, as it will return all columns from the customers table. The second statement is correct, but it is not the only difference between the result sets.

[:Â JOIN | Databricks on AWS,Â JOIN - Azure Databricks - Databricks SQL | Microsoft Learn,Â array_join function | Databricks on AWS,Â Hints | Databricks on AWS, , ]

Question # 5

A data analyst has set up a SQL query to run every four hours on a SQL endpoint, but the SQL endpoint is taking too long to start up with each run.

Which of the following changes can the data analyst make to reduce the start-up time for the endpoint while managing costs?

Reduce the SQL endpoint cluster size

Increase the SQL endpoint cluster size

Turn off the Auto stop feature

Increase the minimum scaling value

Use a Serverless SQL endpoint

Question # 6

What describes Partner Connect in Databricks?

it allows for free use of Databricks partner tools through a common API.

it allows multi-directional connection between Databricks and Databricks partners easier.

It exposes connection information to third-party tools via Databricks partners.

It is a feature that runs Databricks partner tools on a Databricks SQL Warehouse (formerly known as a SQL endpoint).

Question # 7

Which of the following statements describes descriptive statistics?

A branch of statistics that uses summary statistics to quantitatively describe and summarize data.

A branch of statistics that uses a variety of data analysis techniques to infer properties of an underlying distribution of probability.

A branch of statistics that uses quantitative variables that must take on a finite or countably infinite set of values.

A branch of statistics that uses summary statistics to categorically describe and summarize data.

A branch of statistics that uses quantitative variables that must take on an uncountable set of values.

Question # 8

A data analyst has been asked to produce a visualization that shows the flow of users through a website.

Which of the following is used for visualizing this type of flow?

Heatmap

IChoropleth

Word Cloud

Pivot Table

Sankey

Question # 9

How can a data analyst determine if query results were pulled from the cache?

Go to the Query History tab and click on the text of the query. The slideout shows if the results came from the cache.

Go to the Alerts tab and check the Cache Status alert.

Go to the Queries tab and click on Cache Status. The status will be green if the results from the last run came from the cache.

Go to the SQL Warehouse (formerly SQL Endpoints) tab and click on Cache. The Cache file will show the contents of the cache.

Go to the Data tab and click Last Query. The details of the query will show if the results came from the cache.

Question # 10

A data organization has a team of engineers developing data pipelines following the medallion architecture using Delta Live Tables. While the data analysis team working on a project is using gold-layer tables from these pipelines, they need to perform some additional processing of these tables prior to performing their analysis.

Which of the following terms is used to describe this type of work?

Data blending

Last-mile

Data testing

Last-mile ETL

Data enhancement

Winter Sale Limited Time 65% Discount Offer - Ends in 0d 00h 00m 00s - Coupon code: ecus65

Databricks Databricks-Certified-Data-Analyst-Associate - Databricks Certified Data Analyst Associate Exam

The Answer Is:

Explanation:

The Answer Is:

Explanation:

The Answer Is:

Explanation:

The Answer Is:

Explanation:

The Answer Is:

Explanation:

The Answer Is:

Explanation:

The Answer Is:

Explanation:

The Answer Is:

Explanation:

The Answer Is:

Explanation:

The Answer Is:

Explanation: