[Music] [Sameeksha Arora] Hello, everyone. Welcome to today's session. I am Sameeksha Arora, a Senior Product Manager at Adobe Experience Platform. In this session, we will explore key tips and best practices to help you maximize the value of Data Distiller, enabling you to transform data efficiently and derive impactful insights. So let's get started.
To set the stage for today's discussion, let's take a look at these images. A chef and a musician, both masters of their craft, shows that while tools are important, the true excellence comes from expertise and precision. Similarly, Data Distiller provides the tools to refine raw data, but the real value comes from applying strategic thinking to turn data into actionable insights.
Let's take a closer look at what this means in practice.
Just as a chef blends perfectly well ingredients to create a masterpiece, data must be refined to unlock real value. And guess what? When done right, it's not just structured data, it's an orchestrated experience.
Data Distiller enables this transformation, making raw data actionable.
And here's what we will cover today. We'll explore Data Distiller overview, its key capabilities including derived datasets, SQL Insights, and AI/ML feature pipeline, followed by a real world RFM-based SQL Audience use case. And then we'll dive into some of the best practices for efficiency.
Let's begin with an overview of Adobe Experience Platform Data Distiller.
Before diving into the best practices, let's first understand where did our Distiller fits within the Adobe Experience Cloud ecosystem? It's an add-on SKU that enables data curation and transformation, ensuring that data is optimized for downstream applications.
Now let's take a look at its impact.
Data Distiller enables you to explore, prepare, and curate high value data to create datasets that can be directly leveraged by downstream applications like Real-Time CDP, Customer Journey Analytics, or Adobe Journey Optimizer. By automating this dataset creation, Data Distiller not only saves time, but it ensures that data is ready for analysis.
This makes it easier to generate insights, derive AI/ML workflows, ultimately improving the decision making across your business. Now that we know about its role, let's see how Data Distiller turns raw data into refined, actionable insights.
Data Distiller is very carefully positioned after the data ingestion phase, where raw data enters the platform and once ingested, Data Distiller transforms, cleans, and curates this data to make it actionable. It automates the creation of high quality, enriched datasets that are ready for analysis. And from there, the curated data flows seamlessly into downstream applications such as RTCDP, AJO, and CJA, where further insights and actions can be derived. Essentially, Data Distiller acts as a bridge between your raw data and actionable insights, preparing data for advanced analytics and workflows.
Raw data is like unprepared ingredients. It needs some kind of processing. That's why Data Distiller enables you to do data transformation that converts your raw data into structured format that can be easily analyzed. Next, with data processing and enrichment, you define the data by applying filters, transformations, and business rules, ensuring it meets specific business needs. Now what? Once your data is processed, the data seamlessly integrates with downstream tools ready for customer engagement strategies.
And finally, Data Distiller enables AI/ML optimization, allowing organizations to extract maximum value from their processed data for better decision making and personalization.
Now that we know how it helps. Let's see what Data Distiller does, and let's also explore how to use its capabilities effectively.
Data Distiller essentially has three core capabilities. Derived Datasets, it enables you to create customized datasets for targeted analysis, such as audience segmentation or if you want to do churn analysis, you can do all that by creating derived datasets. With SQL Insights, you can use Query Pro Mode to write SQL queries, visualize data, or simply just integrate the data with BI tools like Power BI and Tableau. And lastly, with AI/ML feature pipeline offering, you can enrich ML workflows using curated data, allowing data scientists to compute features, train models, and deploy predictions seamlessly. If you would like to learn more about any of these capabilities, please follow the link shared in Read More section.
Now let's see these capabilities in action with an RFM-based SQL Audience use case. Let's start by taking a look at Luma Store's business.
Luma Store is a leading sports fitness apparel brand that caters to health-conscious individuals and athletes.
The goal is to enhance marketing effectiveness by sub-segmenting customers and personalizing engagement, ultimately improving targeting and driving growth.
Currently, their approach is broad, generic, and lacks personalization. A more refined strategy is needed to improve retention and customer's lifetime value.
You must choose future marketing strategy will focus on enhanced segmentation. Instead of just using demographics, focus on using behavioral and transactional data to understand your customers better. Imagine you are the marketing manager at the Luma Store. What would be your goal? Your goal should be to build a marketing strategy with meaningful customer segments, and for this purpose, you aim to target customers based on their past behavior using RFM segmentation.
RFM stands for Recency, Frequency, and Monetary. It's a data-driven approach to do customer segmentation and analysis.
By leveraging RFM, you can send personalized messages and you can additionally offer tailored rewards, and you can additionally unify omnichannel experiences for your customers.
Let's learn a little bit more about RFM. RFM segmentation classifies customer based on three factors. Recency. Recency gauges the time elapsed since a customer's last purchase, providing insights into engagement levels and future transaction potential. Frequency. Frequency assesses the frequency of customer interactions, serving as an indicator of loyalty and sustained engagement. And lastly, monetary. Monetary measures the total spending of customers emphasizing their value to the business. By using these three factors, it allows you to do more precise targeting.
The combination of these three factors R, F, and M enables businesses to assign a numeric scores to each customer, typically on a scale from one to four, where lower scores signify more favorable outcomes. For instance, a customer scoring one in all categories is deemed to be the champion, showcasing recent activity, high engagement, and substantial spending. Conversely, customers scoring four are at risk because they signal potential churn in the near future. Hence, these three factors helps us to improve marketing personalization.
Building an RFM audience is like crafting a perfect dish. It follows simple fun steps. First, you need to select the ingredients. Second, you need to enhance the flavors by adding herbs, spices, or any of your special technique. Then, you simply measure and mix your ingredients to create the desired flavor. Now you let it cook on low heat to develop those flavors that you are looking for. And lastly, once it's ready, you plate and serve the dish.
Let's apply the same structured cooking approach to building an RFM audience in Data Distiller.
First, you connect to the data lake to identify relevant datasets, just like you select your ingredients. Then you enrich your data, which means you apply filters and do required data transformation, just like you enhance those flavors by adding herbs and spices. Next, you set segmentation criteria and write curated data to the lake, just like mixing and measuring those ingredients.
Then you just simply schedule your batch processing for automated hydration, just like you let it cook at sim. And lastly, you activate your audience just like you plate and serve the dish. By following this structured approach, you ensured that the audience is as defined, as impactful, and as effective as a perfectly crafted meal. So let's go and see these steps in action.
Before cooking, a chef checks their ingredients. Data selection is similar. Accessing the data lake ensures quality datasets for segmentation.
Just like a chef tastes the ingredients before cooking, data validation works the same way. You just need to run a simple Select query, and it will help you to inspect, validate, and analyze data to ensure that it has been accurately translated during the ingestion process.
This process helps you identify any discrepancies, inconsistencies, or missing information in the data even before you commence segmentation. Isn't it amazing? In this example, luma_web_data is the analytical dataset for the Luma Store.
Now that you have the right data, it's time to refine it.
Data Distiller enriches raw data, applying transformations for better segmentation.
Just as a chef balances the flavor, you can fine-tune customer segmentation using RFM metrics. In this step, you analyze the customers, when customers' last purchase, how often they buy, and how much they spend, much like adjusting flavors for one of your signature dish. By dividing RFM scores into quartiles, you can classify audiences based on marketing goals.
Just like a chef categorizes ingredients for different recipes, you classify customers using the NTILE function for RFM model.
NTILE divides data into equal size groups. Here, it segments customer into four quadrants based on their RFM scores, helping identify VIPs or the champions and those who needs the engagement. The top quartile with value one features VIP customers, while the bottom quartile with value four signals those needing reengagement. This helps you to create VIP and reengagement segments easily.
Now that you have the RFM scores, let's enrich customer profile before we dive into audience creation.
Let's define segmentation criteria and write curated data into the profile store, just like precisely measuring ingredients for a recipe. In Data Distiller, you can run a SQL query to create a Profile-enabled dataset and then insert RFM segmentation data into that dataset. Just make sure to mark userId as one of the PRIMARY IDENTITY. This will help you ensure seamless integration with profile and identity stores, which will further help you in future to do activation.
To keep segmentation updated, you can simply schedule queries for automation.
Just like setting a timer in the kitchen ensures perfect cooking. Scheduling queries ensures that profile gets updated automatically. You don't have to worry about it.
To automate SQL execution, simply select a query template, set its execution frequency, and save the schedule. Profiles stay updated at all times and no manual effort is required. Moving on, just like a recipe guides a dish, structuring an audience ensures segmentation works efficiently. Define your key fields and segment customers based on behavioral patterns, ensuring smooth audience activation. Plating and serving completes a meal. Activating and audience completes segmentation. With just a few SQL queries, you can create audience schema and make sure to define key attributes like userId, orders, total revenue value, and whatever you need to know. And then you can simply insert RFM segmented customer data into the audience schema. Remember, the audience created by SQL is automatically registered under Data Distiller origin, making it ready for the activation. Also, when no longer needed, just use the syntax drop audience to remove outdated segments.
Now that the audience is ready to be used in real world marketing campaigns, SQL created audience are automatically registered in Data Distiller Origin and can be activated instantly in Amazon S3, or SFTP, or Azure. No extra setup is needed. It's all ready to go for personalization and targeting.
Now let's shift gears to best practices for maximizing Data Distiller's potential.
The very first thing that you should remember is that choosing the right tool is a key. Query Service Ad-Hoc is best for quick on-the-fly data exploration and validation. On the other hand, and literally the other hand of this lady, Data Distiller is built for more complex, large scale data processing and scheduled transformations. A key tip is to avoid running complex transformations in ad-hoc queries, instead, simply schedule them in Data Distiller. This will help you to optimize performance, and it will ensure efficient use of your resources. Just like a chef, how a chef preps ingredients in advance rather than rushing in the middle of service. Now before we dive into this best practice, let's first define compute hours.
Compute hours measures the query processing time. The larger the query, the more it consumes. Just like a kitchen where smarter prep saves effort and inefficiency leads to wastage. To optimize schedule queries instead of running them in ad-hoc mode, or instead of running them manually, use the SNAPSHOT clause to process only newly ingested data. Avoid re-processing of data unnecessarily. This ensures every single compute hour is used effectively.
And now let's look at some of the more key best practices for optimizing utilization of Data Distiller.
Use Query Quarantine to isolate failing queries, but pair it with other troubleshooting manual methods. Data Distiller today supports PostgreSQL. Python execution is not available.
Be mindful of your data export limits based on your licensing entitlements. Additionally, Data Distiller supports advanced ML through SQL commands. You can create models, evaluate them, and make predictions directly using SQL, making it a more powerful tool beyond just querying. And lastly, check out our Experience League for more tools, guides, and best practices.
Now let's quickly recap.
Before we wrap up, here are the final takeaways.
Data Distiller transforms raw data into actionable insights. RFM analysis helps segment customers for personalized marketing. And lastly, best practices improve efficiency and optimize resource usage. Applying these strategy will maximize Data Distiller's potential and enhance data-driven decision making.
Moving on, I would like to thank you for your time today. With Data Distiller, you now have the tools to turn data into actionable insights and derive impactful decisions. To help you get started right away, the first link provided here helps you to execute the use case that we just discussed today. Additionally, to get you started, it provides you with a sample data in CSV format followed by SQL queries to run the real-world use case that we just discussed. And lastly, it gives you a step-by-step guide to help you execute. So just dive in and explore. Thank you. [Music]