Full Professional-Data-Engineer Practice Test and 165 unique questions with explanations waiting just for you!
Google Cloud Certified Dumps Professional-Data-Engineer Exam for Full Questions - Exam Study Guide
NEW QUESTION 39
What are all of the BigQuery operations that Google charges for?
- A. Storage, queries, and streaming inserts
- B. Storage, queries, and exporting data
- C. Storage, queries, and loading data from a file
- D. Queries and streaming inserts
Answer: A
Explanation:
Explanation
Google charges for storage, queries, and streaming inserts. Loading data from a file and exporting data are free operations.
Reference: https://cloud.google.com/bigquery/pricing
NEW QUESTION 40
Your neural network model is taking days to train. You want to increase the training speed. What can you do?
- A. Subsample your test dataset.
- B. Subsample your training dataset.
- C. Increase the number of input features to your model.
- D. Increase the number of layers in your neural network.
Answer: D
Explanation:
Explanation/Reference: https://towardsdatascience.com/how-to-increase-the-accuracy-of-a-neural-network-9f5d1c6f407d
NEW QUESTION 41
You are using Google BigQuery as your data warehouse. Your users report that the following simple query is running very slowly, no matter when they run the query:
SELECT country, state, city FROM [myproject:mydataset.mytable] GROUP BY country You check the query plan for the query and see the following output in the Read section of Stage:1:
What is the most likely cause of the delay for this query?
- A. Users are running too many concurrent queries in the system
- B. Either the state or the city columns in the [myproject:mydataset.mytable] table have too many NULL values
- C. The [myproject:mydataset.mytable] table has too many partitions
- D. Most rows in the [myproject:mydataset.mytable] table have the same value in the country column, causing data skew
Answer: D
NEW QUESTION 42
Your company needs to upload their historic data to Cloud Storage. The security rules don't allow access from external IPs to their on-premises resources. After an initial upload, they will add new data from existing on-premises applications every day. What should they do?
- A. Install an FTP server on a Compute Engine VM to receive the files and move them to Cloud Storage.
- B. Use Cloud Dataflow and write the data to Cloud Storage.
- C. Execute gsutil rsync from the on-premises servers.
- D. Write a job template in Cloud Dataproc to perform the data transfer.
Answer: C
NEW QUESTION 43
You are building a new application that you need to collect data from in a scalable way. Data arrives continuously from the application throughout the day, and you expect to generate approximately 150 GB of JSON data per day by the end of the year. Your requirements are:
* Decoupling producer from consumer
* Space and cost-efficient storage of the raw ingested data, which is to be stored indefinitely
* Near real-time SQL query
* Maintain at least 2 years of historical data, which will be queried with SQL Which pipeline should you use to meet these requirements?
- A. Create an application that publishes events to Cloud Pub/Sub, and create a Cloud Dataflow pipeline that transforms the JSON event payloads to Avro, writing the data to Cloud Storage and BigQuery.
- B. Create an application that writes to a Cloud SQL database to store the data. Set up periodic exports of the database to write to Cloud Storage and load into BigQuery.
- C. Create an application that publishes events to Cloud Pub/Sub, and create Spark jobs on Cloud Dataproc to convert the JSON data to Avro format, stored on HDFS on Persistent Disk.
- D. Create an application that provides an API. Write a tool to poll the API and write data to Cloud Storage as gzipped JSON files.
Answer: D
NEW QUESTION 44
A shipping company has live package-tracking data that is sent to an Apache Kafka stream in real time. This is then loaded into BigQuery. Analysts in your company want to query the tracking data in BigQuery to analyze geospatial trends in the lifecycle of a package. The table was originally created with ingest-date partitioning.
Over time, the query processing time has increased. You need to implement a change that would improve query performance in BigQuery. What should you do?
- A. Implement clustering in BigQuery on the ingest date column.
- B. Implement clustering in BigQuery on the package-tracking ID column.
- C. Re-create the table using data partitioning on the package delivery date.
- D. Tier older data onto Google Cloud Storage files and create a BigQuery table using GCS as an external data source.
Answer: B
NEW QUESTION 45
You are using Google BigQuery as your data warehouse. Your users report that the following simple query is running very slowly, no matter when they run the query:
SELECT country, state, city FROM [myproject:mydataset.mytable] GROUP BY country You check the query plan for the query and see the following output in the Read section of Stage:1:
What is the most likely cause of the delay for this query?
- A. Users are running too many concurrent queries in the system
- B. Most rows in the [myproject:mydataset.mytable]table have the same value in the country column, causing data skew
- C. Either the state or the city columns in the [myproject:mydataset.mytable]table have too many NULL values
- D. The [myproject:mydataset.mytable] table has too many partitions
Answer: A
NEW QUESTION 46
You are building new real-time data warehouse for your company and will use Google BigQuery streaming inserts. There is no guarantee that data will only be sent in once but you do have a unique ID for each row of data and an event timestamp. You want to ensure that duplicates are not included while interactively querying data. Which query type should you use?
- A. Include ORDER BY DESK on timestamp column and LIMIT to 1.
- B. Use the ROW_NUMBER window function with PARTITION by unique ID along with WHERE row equals 1.
- C. Use GROUP BY on the unique ID column and timestamp column and SUM on the values.
- D. Use the LAG window function with PARTITION by unique ID along with WHERE LAG IS NOT NULL.
Answer: B
NEW QUESTION 47
Suppose you have a dataset of images that are each labeled as to whether or not they contain a human face. To create a neural network that recognizes human faces in images using this labeled dataset, what approach would likely be the most effective?
- A. Use K-means Clustering to detect faces in the pixels.
- B. Use deep learning by creating a neural network with multiple hidden layers to automatically detect features of faces.
- C. Use feature engineering to add features for eyes, noses, and mouths to the input data.
- D. Build a neural network with an input layer of pixels, a hidden layer, and an output layer with two categories.
Answer: B
Explanation:
Traditional machine learning relies on shallow nets, composed of one input and one output layer, and at most one hidden layer in between. More than three layers (including input and output) qualifies as "deep" learning. So deep is a strictly defined, technical term that means more than one hidden layer.
In deep-learning networks, each layer of nodes trains on a distinct set of features based on the previous layer's output. The further you advance into the neural net, the more complex the features your nodes can recognize, since they aggregate and recombine features from the previous layer.
A neural network with only one hidden layer would be unable to automatically recognize high-level features of faces, such as eyes, because it wouldn't be able to "build" these features using previous hidden layers that detect low-level features, such as lines. Feature engineering is difficult to perform on raw image data.
K-means Clustering is an unsupervised learning method used to categorize unlabeled data.
Reference: https://deeplearning4j.org/neuralnet-overview
NEW QUESTION 48
You designed a database for patient records as a pilot project to cover a few hundred patients in three clinics. Your design used a single database table to represent all patients and their visits, and you used self-joins to generate reports. The server resource utilization was at 50%. Since then, the scope of the project has expanded. The database must now store 100 times more patient records. You can no longer run the reports, because they either take too long or they encounter errors with insufficient compute resources. How should you adjust the database design?
- A. Add capacity (memory and disk space) to the database server by the order of 200.
- B. Partition the table into smaller tables, with one for each clinic. Run queries against the smaller table pairs, and use unions for consolidated reports.
- C. Shard the tables into smaller ones based on date ranges, and only generate reports with prespecified date ranges.
- D. Normalize the master patient-record table into the patient table and the visits table, and create other necessary tables to avoid self-join.
Answer: C
NEW QUESTION 49
You have some data, which is shown in the graphic below. The two dimensions are X and Y, and the
shade of each dot represents what class it is. You want to classify this data accurately using a linear
algorithm. To do this you need to add a synthetic feature. What should the value of that feature be?
- A. Y^2
- B. X^2
- C. cos(X)
- D. X^2+Y^2
Answer: C
NEW QUESTION 50
You are operating a streaming Cloud Dataflow pipeline. Your engineers have a new version of the pipeline with a different windowing algorithm and triggering strategy. You want to update the running pipeline with the new version. You want to ensure that no data is lost during the update. What should you do?
- A. Update the Cloud Dataflow pipeline inflight by passing the --update option with the --jobName set to the existing job name
- B. Stop the Cloud Dataflow pipeline with the Drain option. Create a new Cloud Dataflow job with the updated code
- C. Stop the Cloud Dataflow pipeline with the Cancel option. Create a new Cloud Dataflow job with the updated code
- D. Update the Cloud Dataflow pipeline inflight by passing the --update option with the --jobName set to a new unique job name
Answer: A
NEW QUESTION 51
You plan to deploy Cloud SQL using MySQL. You need to ensure high availability in the event of a zone failure.
What should you do?
- A. Create a Cloud SQL instance in one zone, and configure an external read replica in a zone in a different region.
- B. Create a Cloud SQL instance in one zone, and create a failover replica in another zone within the same region.
- C. Create a Cloud SQL instance in one zone, and create a read replica in another zone within the same region.
- D. Create a Cloud SQL instance in a region, and configure automatic backup to a Cloud Storage bucket in the same region.
Answer: A
NEW QUESTION 52
You decided to use Cloud Datastore to ingest vehicle telemetry data in real time. You want to build a storage system that will account for the long-term data growth, while keeping the costs low. You also want to create snapshots of the data periodically, so that you can make a point-in-time (PIT) recovery, or clone a copy of the data for Cloud Datastore in a different environment. You want to archive these snapshots for a long time. Which two methods can accomplish this? Choose 2 answers.
- A. Use managed exportm, and then import to Cloud Datastore in a separate project under a unique namespace reserved for that export.
- B. Write an application that uses Cloud Datastore client libraries to read all the entities. Format the exported data into a JSON file. Apply compression before storing the data in Cloud Source Repositories.
- C. Use managed export, and then import the data into a BigQuery table created just for that export, and delete temporary export files.
- D. Write an application that uses Cloud Datastore client libraries to read all the entities. Treat each entity as a BigQuery table row via BigQuery streaming insert. Assign an export timestamp for each export, and attach it as an extra column for each row. Make sure that the BigQuery table is partitioned using the export timestamp column.
- E. Use managed export, and store the data in a Cloud Storage bucket using Nearline or Coldline class.
Answer: B,C
NEW QUESTION 53
Flowlogistic Case Study
Company Overview
Flowlogistic is a leading logistics and supply chain provider. They help businesses throughout the world
manage their resources and transport them to their final destination. The company has grown rapidly,
expanding their offerings to include rail, truck, aircraft, and oceanic shipping.
Company Background
The company started as a regional trucking company, and then expanded into other logistics market.
Because they have not updated their infrastructure, managing and tracking orders and shipments has
become a bottleneck. To improve operations, Flowlogistic developed proprietary technology for tracking
shipments in real time at the parcel level. However, they are unable to deploy it because their technology
stack, based on Apache Kafka, cannot support the processing volume. In addition, Flowlogistic wants to
further analyze their orders and shipments to determine how best to deploy their resources.
Solution Concept
Flowlogistic wants to implement two concepts using the cloud:
Use their proprietary technology in a real-time inventory-tracking system that indicates the location of
their loads
Perform analytics on all their orders and shipment logs, which contain both structured and unstructured
data, to determine how best to deploy resources, which markets to expand info. They also want to use
predictive analytics to learn earlier when a shipment will be delayed.
Existing Technical Environment
Flowlogistic architecture resides in a single data center:
Databases
8 physical servers in 2 clusters
- SQL Server - user data, inventory, static data
3 physical servers
- Cassandra - metadata, tracking messages
10 Kafka servers - tracking message aggregation and batch insert
Application servers - customer front end, middleware for order/customs
60 virtual machines across 20 physical servers
- Tomcat - Java services
- Nginx - static content
- Batch servers
Storage appliances
- iSCSI for virtual machine (VM) hosts
- Fibre Channel storage area network (FC SAN) - SQL server storage
- Network-attached storage (NAS) image storage, logs, backups
10 Apache Hadoop /Spark servers
- Core Data Lake
- Data analysis workloads
20 miscellaneous servers
- Jenkins, monitoring, bastion hosts,
Business Requirements
Build a reliable and reproducible environment with scaled panty of production.
Aggregate data in a centralized Data Lake for analysis
Use historical data to perform predictive analytics on future shipments
Accurately track every shipment worldwide using proprietary technology
Improve business agility and speed of innovation through rapid provisioning of new resources
Analyze and optimize architecture for performance in the cloud
Migrate fully to the cloud if all other requirements are met
Technical Requirements
Handle both streaming and batch data
Migrate existing Hadoop workloads
Ensure architecture is scalable and elastic to meet the changing demands of the company.
Use managed services whenever possible
Encrypt data flight and at rest
Connect a VPN between the production data center and cloud environment
SEO Statement
We have grown so quickly that our inability to upgrade our infrastructure is really hampering further growth
and efficiency. We are efficient at moving shipments around the world, but we are inefficient at moving
data around.
We need to organize our information so we can more easily understand where our customers are and
what they are shipping.
CTO Statement
IT has never been a priority for us, so as our data has grown, we have not invested enough in our
technology. I have a good staff to manage IT, but they are so busy managing our infrastructure that I
cannot get them to do the things that really matter, such as organizing our data, building the analytics, and
figuring out how to implement the CFO' s tracking technology.
CFO Statement
Part of our competitive advantage is that we penalize ourselves for late shipments and deliveries. Knowing
where out shipments are at all times has a direct correlation to our bottom line and profitability.
Additionally, I don't want to commit capital to building out a server environment.
Flowlogistic's CEO wants to gain rapid insight into their customer base so his sales team can be better
informed in the field. This team is not very technical, so they've purchased a visualization tool to simplify
the creation of BigQuery reports. However, they've been overwhelmed by all the data in the table, and are
spending a lot of money on queries trying to find the data they need. You want to solve their problem in the
most cost-effective way. What should you do?
- A. Export the data into a Google Sheet for virtualization.
- B. Create an additional table with only the necessary columns.
- C. Create identity and access management (IAM) roles on the appropriate columns, so only they appear
in a query. - D. Create a view on the table to present to the virtualization tool.
Answer: D
NEW QUESTION 54
What are two methods that can be used to denormalize tables in BigQuery?
- A. 1) Join tables into one table; 2) Use nested repeated fields
- B. 1) Use nested repeated fields; 2) Use a partitioned table
- C. 1) Split table into multiple tables; 2) Use a partitioned table
- D. 1) Use a partitioned table; 2) Join tables into one table
Answer: A
Explanation:
The conventional method of denormalizing data involves simply writing a fact, along with all its dimensions, into a flat table structure. For example, if you are dealing with sales transactions, you would write each individual fact to a record, along with the accompanying dimensions such as order and customer information.
The other method for denormalizing data takes advantage of BigQuery's native support for nested and repeated structures in JSON or Avro input data. Expressing records using nested and repeated structures can provide a more natural representation of the underlying data. In the case of the sales order, the outer part of a JSON structure would contain the order and customer information, and the inner part of the structure would contain the individual line items of the order, which would be represented as nested, repeated elements.
Reference: https://cloud.google.com/solutions/bigquery-data-
warehouse#denormalizing_data
NEW QUESTION 55
Which of these are examples of a value in a sparse vector? (Select 2 answers.)
- A. [0, 1]
- B. [0, 5, 0, 0, 0, 0]
- C. [0, 0, 0, 1, 0, 0, 1]
- D. [1, 0, 0, 0, 0, 0, 0]
Answer: A,D
Explanation:
Categorical features in linear models are typically translated into a sparse vector in which each possible value has a corresponding index or id. For example, if there are only three possible eye colors you can represent 'eye_color' as a length 3 vector: 'brown' would become [1, 0, 0], 'blue' would become [0, 1, 0] and 'green' would become [0, 0, 1]. These vectors are called "sparse" because they may be very long, with many zeros, when the set of possible values is very large (such as all English words).
[0, 0, 0, 1, 0, 0, 1] is not a sparse vector because it has two 1s in it. A sparse vector contains only a single
1.
[0, 5, 0, 0, 0, 0] is not a sparse vector because it has a 5 in it. Sparse vectors only contain 0s and 1s.
Reference: https://www.tensorflow.org/tutorials/linear#feature_columns_and_transformations
NEW QUESTION 56
You want to automate execution of a multi-step data pipeline running on Google Cloud. The pipeline includes Cloud Dataproc and Cloud Dataflow jobs that have multiple dependencies on each other. You want to use managed services where possible, and the pipeline will run every day. Which tool should you use?
- A. Workflow Templates on Cloud Dataproc
- B. Cloud Scheduler
- C. cron
- D. Cloud Composer
Answer: D
NEW QUESTION 57
When running a pipeline that has a BigQuery source, on your local machine, you continue to get permission denied errors. What could be the reason for that?
- A. Pipelines cannot be run locally
- B. BigQuery cannot be accessed from local machines
- C. Your gcloud does not have access to the BigQuery resources
- D. You are missing gcloud on your machine
Answer: C
Explanation:
Explanation
When reading from a Dataflow source or writing to a Dataflow sink using DirectPipelineRunner, the Cloud Platform account that you configured with the gcloud executable will need access to the corresponding source/sink Reference:
https://cloud.google.com/dataflow/java-sdk/JavaDoc/com/google/cloud/dataflow/sdk/runners/DirectPipelineRun
NEW QUESTION 58
What are all of the BigQuery operations that Google charges for?
- A. Storage, queries, and streaming inserts
- B. Storage, queries, and exporting data
- C. Storage, queries, and loading data from a file
- D. Queries and streaming inserts
Answer: A
Explanation:
Google charges for storage, queries, and streaming inserts. Loading data from a file and exporting data are free operations.
Reference: https://cloud.google.com/bigquery/pricing
NEW QUESTION 59
You are creating a new pipeline in Google Cloud to stream IoT data from Cloud Pub/Sub through Cloud Dataflow to BigQuery. While previewing the data, you notice that roughly 2% of the data appears to be corrupt. You need to modify the Cloud Dataflow pipeline to filter out this corrupt data. What should you do?
- A. Add a ParDo transform in Cloud Dataflow to discard corrupt elements.
- B. Add a Partition transform in Cloud Dataflow to separate valid data from corrupt data.
- C. Add a SideInput that returns a Boolean if the element is corrupt.
- D. Add a GroupByKey transform in Cloud Dataflow to group all of the valid data together and discard the rest.
Answer: A
NEW QUESTION 60
You work for a large fast food restaurant chain with over 400,000 employees. You store employee information in Google BigQuery in a Userstable consisting of a FirstNamefield and a LastNamefield. A member of IT is building an application and asks you to modify the schema and data in BigQuery so the application can query a FullNamefield consisting of the value of the FirstNamefield concatenated with a space, followed by the value of the LastNamefield for each employee. How can you make that data available while minimizing cost?
- A. Add a new column called FullNameto the Users table. Run an UPDATEstatement that updates the FullNamecolumn for each user with the concatenation of the FirstNameand LastNamevalues.
- B. Create a view in BigQuery that concatenates the FirstNameand LastNamefield values to produce the FullName.
- C. Use BigQuery to export the data for the table to a CSV file. Create a Google Cloud Dataproc job to process the CSV file and output a new CSV file containing the proper values for FirstName, LastNameand FullName. Run a BigQuery load job to load the new CSV file into BigQuery.
- D. Create a Google Cloud Dataflow job that queries BigQuery for the entire Userstable, concatenates the FirstNamevalue and LastNamevalue for each user, and loads the proper values for FirstName, LastName, and FullNameinto a new table in BigQuery.
Answer: D
NEW QUESTION 61
You want to use a database of information about tissue samples to classify future tissue samples as either normal or mutated. You are evaluating an unsupervised anomaly detection method for classifying the tissue samples. Which two characteristic support this method? (Choose two.)
- A. You expect future mutations to have similar features to the mutated samples in the database.
- B. There are very few occurrences of mutations relative to normal samples.
- C. You already have labels for which samples are mutated and which are normal in the database.
- D. There are roughly equal occurrences of both normal and mutated samples in the database.
- E. You expect future mutations to have different features from the mutated samples in the database.
Answer: D,E
Explanation:
Explanation/Reference:
NEW QUESTION 62
......
Authentic Best resources for Professional-Data-Engineer Online Practice Exam: https://www.test4engine.com/Professional-Data-Engineer_exam-latest-braindumps.html
Get the superior quality Professional-Data-Engineer Dumps Questions from Test4Engine: https://drive.google.com/open?id=13mLZucVnfImfUfEavZWHpCNcdUhRj9kE