Skip to main content
In Designer, you can sample the data output by many components in your pipelines. Sampling the data lets you see the data that will be passed to downstream components once the component has performed its designated tasks. Sampling lets you check that the component is configured correctly before running the pipeline. You can also view metadata about the output of many components in Designer. This metadata provides information about the columns that will be passed to downstream components in your pipeline.

Metadata

To view the metadata about the output of a component, select the component on the canvas and click the Metadata tab at the bottom of the canvas. The Metadata tab contains a read-only table showing the name, data type, and size of each column in the component’s output. Use the Text mode toggle to switch between the table view and text mode. You can’t edit the metadata shown in text mode, but you can copy it to use it as the basis for creating a table using the Create Table component. To change the data type of a column in your component’s output, use the Convert Type transformation component.

Sampling data

Before you can sample a component’s output, you must first validate the pipeline by clicking Validate in the top right of the Designer canvas.
To sample the data output by a component, select the component on the canvas and either:
  • Click the Sample data tab at the bottom of the canvas, then click Sample.
  • Click Sample data in the top right of the component properties panel.
Maia will start sampling the component’s output, as shown by the spinning border around the component. If you leave the Sample data tab while sampling is in progress, a notification will appear when the sample is ready—click View in the notification to open the Sample data tab and view the component’s output. The Sample data tab shows the name of the sampled component, a table containing the component’s output, and the number of sampled rows and total rows in the component’s output. For example, Total rows: 25 (120) means that you are currently viewing a sample of 25 out of 120 rows of data. If you close the Sample data tab and return, the sampled data and any applied filters remain available.
Up to 5 MB of sampled data remains available in the Sample data tab across multiple components. If sampling the output of another component would exceed the 5 MB limit, the oldest stored sample will be cleared so that you can sample the current component.For example, if you sample the output of Component A, then Component B, then Component C, and each sample is 1.5 MB, the total sampled data is 4.5 MB, so all three samples will remain available. If you then sample the output of Component D and this sample is 1.5 MB, this would exceed the 5 MB limit. As a result, the sampled output from Component A is cleared so that you can sample Component D.
In the Sample data tab, you can perform the following actions:
  • Sample more data: Use the drop-down menu in the top right, then click Sample to change how many rows of data are sampled. The default is 25 rows, with options between 1 and 1000 rows available.
  • Filter sampled data: Click Filter with Maia to ask Maia Team to help you filter the sample, or enter a filter expression in the Filter field. For more information, read Filtering sampled data.
  • Sort sampled data: Click the header of any column to sort the sampled data by the values in this column. Click once to sort in ascending order, twice to sort in descending order, and three times to remove sorting. The arrow icons in each column header indicate which column is being used to sort the data, and in what order.
  • Export sampled data: Click the Download CSV icon in the bottom right to export a CSV file containing the data currently shown in the sample table.
  • Resize columns: Click and drag the border of a column to change its width.
  • Cleanse data: Click Cleanse data to add a Data Cleanse component to your pipeline, immediately downstream of the component you’re sampling.

Filtering sampled data

You can filter the rows displayed in the table in the Sample data tab using one or more filter queries. Maia sends filter queries to your warehouse as an SQL WHERE clause, and returns data that matches the filter conditions. To filter the data shown in the Sample data table, either:
  • Click Ask Maia next to the Filter field and tell Maia Team how you want to filter the sampled data. Maia Team will create and apply a filter.
  • Enter a filter query in the Filter field at the top of the Sample data tab, then press Enter.
The filters currently applied to the sample are shown above the table, where you can:
  • Remove individual filters: Click the x icon next to a filter to remove it. This deletes the filter and refreshes the sample.
  • Enable or disable all filters: Toggle Filters on to apply all your current filters. Toggle Filters off to disable all current filters without deleting them. Changing this toggle refreshes the sample.

Writing filter queries

When writing your filter query, follow these rules:
  • For Snowflake, enter column names either in UPPERCASE, e.g. ORDER_QUANTITY, or surrounded by double quotes, e.g. "order_quantity".
  • For Amazon Redshift and Databricks, enter column names without any quote marks, e.g. order_quantity.
  • Use single quotes around string values.
  • Do not use quotes around number values.
  • Enter date values in the format YYYY-MM-DD surrounded by single quotes, e.g. '2024-12-31'.
The following examples show how you can create filter queries for different value types and combine query clauses. These filters are written for Snowflake, which is why the column names are in double quotes.
  • Filter for string values: "customer_surname" = 'Smith' will display rows where the customer’s surname is “Smith”.
  • Filter for number values: "order_quantity" > 5 will display rows where the order quantity is greater than 5.
  • Filter for date values: "order_date" < '2025-01-01' will display rows where the order date is before January 1, 2025.
You can combine query clauses using AND and OR, and search for the opposite of a condition using NOT.
  • Using AND: "customer_organization" = 'Matillion' AND "order_date" > '2025-04-01' will display rows for all orders placed by Matillion after April 1, 2025. The data displayed must meet both of these conditions.
  • Using OR: "country" = 'UK' OR "customer_organization" = 'Matillion' will display rows for all orders placed in the UK, regardless of the customer’s organization, and all orders placed by Matillion, regardless of the country. The data displayed only needs to meet one of the conditions.
  • Using NOT: NOT "country" = 'UK' will display rows for all orders placed in a country other than the UK.

Operators

You can use the following operators in your query:
  • Comparison operators
    • =: equals
    • != or <>: does not equal
    • < and >: less than, greater than
    • <= and >=: less than or equal to, greater than or equal to
    • IS NULL: is a null value
    • IS NOT NULL: is a non-null value
  • Logical operators
    • AND: filter results that meet more than one condition
    • OR: filter results that meet at least one of a number of conditions
    • NOT: filter results that meet the opposite of a condition
  • Set operators
    • IN: matches a value within a list or subquery
    • NOT IN: does not match a value within a list or subquery
    • BETWEEN: is within a specified range
    • NOT BETWEEN: is not within a specified range
    • LIKE: matches a pattern (case-sensitive)
    • NOT LIKE: does not match a pattern (case-sensitive)
    • ILIKE: matches a pattern (case-insensitive)
  • Other operators:
    • ||: string concatenation

Warehouse-specific operators

Additionally, there are some operators available for use with each cloud data warehouse.

Snowflake

  • Comparison and logical operators
    • <=>: NULL-safe equals (useful for comparing NULL values)
    • RLIKE: matches a regular expression
    • NOT RLIKE: does not match a regular expression
    • NOT ILIKE: does not match a pattern (case-insensitive)
  • Bitwise operators
    • &: Bitwise AND
    • |: Bitwise OR
    • ^: Bitwise XOR
    • ~: Bitwise NOT

Amazon Redshift

  • Comparison and logical operators
    • SIMILAR TO: matches a specified pattern with SQL regular expressions
    • NOT SIMILAR TO: does not match a specified pattern with SQL regular expressions
  • Arithmetic operators
    • %: modulo

Databricks

  • Comparison and logical operators
    • RLIKE: matches a regular expression
    • NOT RLIKE: does not match a regular expression
    • DIV: integer division
  • Arithmetic operators
    • MOD or %: modulo

Enabling and disabling sampling for a project

If necessary, you can enable and disable sampling at the project level. This is useful if your data contains personal information that cannot be viewed outside your region. Only users with project admin permissions can change this setting. To enable or disable sampling for a project:
  1. In the Your projects page, click the three dots … next to the intended project.
  2. Click Enable sampling or Disable sampling as required.