> For the complete documentation index, see [llms.txt](https://docs.ibexa.ai/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.ibexa.ai/tutorials/building-an-data-extraction-agent.md).

# Building a Data Extraction Agent

| **Time to complete** | \~25 minutes                                                                                                                                      |
| -------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Who this is for**  | Marketing team members, Content Editors, and anyone who needs to collect and organise data from websites - without writing a single line of code. |

## What You Will Build

By the end of this tutorial you will have a working **Product Data Extraction Agent** - an agent that visits a web page, reads the product listings on it, and returns a clean, structured list of products (names, prices, and any other details it finds). You will use **Tavily** - a web search and content extraction service - to give the agent the ability to read live web pages. Once the agent is set up, you can point it at any publicly accessible product catalogue and have it pull data back for you in seconds. In this tutorial the agent will extract the product catalogue from <https://sandbox.oxylabs.io/products> - a public web scraping sandbox that contains a fictional games store, specifically provided for practising data collection tasks like this one.

## Prerequisites

Before you start, make sure the following are in place:

* At least one **AI model** has been connected and enabled. Go to **Organisation → AI Models** and confirm a model is listed. If none are available, ask your administrator
* You can create a free Tavily account - instructions are in Step 1 below.

## Overview of the Steps

1. Create a Tavily account and get your API key
2. Add Tavily as an MCP Server in the platform
3. Create the data extraction agent
4. Run the agent and review the results

## Step 1: Create a Tavily Account and Get Your API Key

**Tavily** (<https://www.tavily.com/>) is a service that lets AI agents search the web and extract content from specific web pages. In this tutorial you will use its **extract** capability to pull product data from a page. Tavily offers a free tier that is more than enough to complete this tutorial and run the agent regularly.

### 1.1 - Sign up for Tavily

1. Open <https://www.tavily.com/> in your browser.
2. Select **Try it for free**
3. Create an account using your email address, or sign in with the identity provider of your choice.
4. Confirm your email address if prompted.

### 1.2 - Get your API key

Once you are logged in to the Tavily dashboard:

1. Go to the **API Keys** section (listed in the left sidebar of the dashboard).
2. Your default API key is shown on this page. It looks like: `tvly-xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx`.
3. Select the **copy** icon next to the key to copy it to your clipboard. **Keep this key safe.** Treat it like a password - do not share it publicly or paste it into a chat or email. You will add it to the platform in the next step using a secure header field. **Free tier limits**: Tavily's free plan includes 1,000 API credits per month. Each page extraction uses a small number of credits. For occasional use this tutorial will cost a fraction of your monthly allowance.

## Step 2: Add Tavily as an MCP Server

Now you will connect Tavily to the platform so that your agent can use it as a tool. MCP Servers are how the platform connects to external services like Tavily.

### 2.1 - Open the MCP Servers page

Go to **Organisation → MCP Servers** in the main navigation menu.

### 2.2 - Add a new server

Select **Add MCP Server** in the top-right corner and find Tavily in MCP Catalog. Fill in the form with the following values:

| Field          | Value                                                     |
| -------------- | --------------------------------------------------------- |
| **Name**       | `Tavily`                                                  |
| **Identifier** | `tavily` *(auto-filled from the name - leave as is)*      |
| **URL**        | `https://mcp.tavily.com/mcp/?tavilyApiKey={YOUR_API_KEY}` |
| **Enabled**    | ✅ *(checked by default - leave it on)*                    |

Replace `{YOUR_API_KEY}` with API Key created in Step 1.2

### 2.3 - Save and verify the server

1. Select **Save**. You are taken to the Tavily server detail page.
2. Open the **actions menu** (top-right of the page) and select **Verify**.
3. Wait a few seconds. A success message confirms the server is reachable and its tools are listed. If verification fails, double-check that your API key is correct and that the `Bearer` prefix (with a space after it) is included in the header value.

> **Tools you will see:** After a successful verification, the Tools Browser on the server detail page shows all tools Tavily provides. The one you will use in this tutorial is **tavily-extract** - it reads the full content of a given URL and returns it in a structured format the agent can work with.

## Step 3: Create the Data Extraction Agent

With Tavily connected, you can now build the agent that will use it to extract product data.

### 3.1 - Open the Agents page

Select **Agents** in the main navigation menu, then select **Add Agent** in the top-right corner. The **Agents Browser** opens. Select **Add Custom Agent** to go directly to the creation wizard.

### 3.2 - Step 1 of 6: Instructions

The instructions are the agent's job description. They tell it exactly what to do, how to format the output, and how to behave if something goes wrong. Paste the following into the **Instructions** field:

```
You are a web data extraction agent. Your job is to extract a structured list of products from a given web page.

When asked to extract products from a URL, follow these steps:
1. Use the tavily-extract tool to fetch the content of the provided URL.
2. Parse the extracted content and identify all products listed on the page.
3. For each product, collect the following details if available:
   - Product name
   - Price
   - Category or genre (if shown)
   - Rating or review score (if shown)
   - Any other relevant attributes visible on the page

4. Return the results as a clearly formatted list. Use the following structure for each product:

**Product Name**
- Price: [price]
- Category: [category]
- Rating: [rating]
- Notes: [any other relevant details]

5. At the end of the list, include a short summary stating:
   - How many products were found in total
   - The price range (lowest and highest price)

If the page cannot be loaded or no products are found, explain what happened clearly and suggest that the user double-checks the URL.

Do not add commentary, opinions, or information that is not present on the page. Report only what you observe.
```

> **Why these instructions?** Clear, step-by-step instructions help the agent stay focused and produce consistently structured output. The formatting rules at the end make the results easy to read and copy into a spreadsheet or report.

Select **Next**.

### 3.3 - Step 2 of 6: Triggers

Triggers define when the agent runs. For this tutorial you want to be able to start it manually on demand.

1. Select **Add trigger**.
2. Choose **Run Now** from the trigger type list.
3. Leave all settings at their defaults. Select **Next**.

### 3.4 - Step 3 of 6: Tools

Tools give the agent the ability to call external services. You need to connect the Tavily server you added in Step 2 and enable the extraction tool.

1. Select **Select tools**.
2. Find **Tavily** in the server list.
3. In the tool list, tick **tavily-extract** - this is the tool that fetches and reads web page content.
4. Confirm your selection and return to the wizard.

> **Why only tavily-extract and not tavily-search?** The `tavily-extract` tool visits a specific URL you provide and returns its content. The `tavily-search` tool performs a web search by keyword. Since you already know the exact URL of the page to extract, `tavily-extract` is the right choice here.

Select **Next**.

### 3.5 - Step 4 of 6: Knowledge Base

| Field      | What to select                                                         |
| ---------- | ---------------------------------------------------------------------- |
| **Access** | **None** - this agent reads from the web, not from internal documents. |

Select **Next**.

### 3.6 - Step 5 of 6: Quality Metrics

Leave this step empty for now - you can configure quality monitoring later from the agent's **Edit** page. Select **Next**.

### 3.7 - Step 6 of 6: General Properties

Fill in the agent's identity, model, and limits:

| Field                       | What to enter                                                                                           |
| --------------------------- | ------------------------------------------------------------------------------------------------------- |
| **Name**                    | `Product Data Extractor`                                                                                |
| **Description**             | `Extracts a structured product list from a given web page URL using Tavily.`                            |
| **Model**                   | Choose the AI model your organisation uses (e.g. *GPT-5.4*). If you are unsure, ask your administrator. |
| **Maximum number of steps** | Leave at the default (**25**) - the extraction task is straightforward and won't need many steps.       |

Leave all other fields at their defaults. Select **Submit**. You are taken to the new agent's detail page.

## Step 4: Run the Agent and Review the Results

### 4.1 - Start a test run

From the agent's **detail page**, open the actions menu (top-right corner) and select **Run test**. A chat interface opens.

### 4.2 - Give the agent its task

Type the following message and press **Enter**:

```
Please extract all products from this page: <https://sandbox.oxylabs.io/products>
```

**About the sandbox URL**: <https://sandbox.oxylabs.io/products> is a publicly available web scraping practice environment maintained by Oxylabs. It contains a fictional games store with product listings (titles, prices, ratings, and genres). The site is provided specifically for practising data extraction - you have full permission to scrape it.

### 4.3 - Wait for the results

The agent will:

1. Call the **tavily-extract** tool with the URL you provided.
2. Receive the page content back from Tavily.
3. Parse the content and format the product list according to its instructions. This typically takes between 15 and 30 seconds. You can watch the agent's progress in the chat interface.

### 4.4 - Review the output

Once complete, the agent returns a structured list of all products found on the page, followed by a summary. It should look something like this:

```
**Grand Theft Auto V**
- Price: $19.99
- Category: Action
- Rating: 4.5 / 5
- Notes: Available in Standard and Premium editions

**The Witcher 3: Wild Hunt**
- Price: $29.99
- Category: RPG
- Rating: 4.8 / 5

...

---
Summary: 16 products found. Price range: $9.99 – $59.99.
```

> **Tip:** You can copy the agent's output and paste it directly into a spreadsheet. Most spreadsheet tools (Google Sheets, Excel) will paste structured text cleanly into rows if you use **Paste special → Paste as plain text**.

### 4.5 - What to do if something goes wrong

| **Symptom**                                | **Likely cause**                                     | **Fix**                                                                                                  |
| ------------------------------------------ | ---------------------------------------------------- | -------------------------------------------------------------------------------------------------------- |
| Agent says it cannot reach the page        | Tavily API key is missing or incorrect               | Go to **Organisation → MCP Servers → Tavily → Edit** and double-check the `Authorization` header value.  |
| Agent returns no products                  | Page structure changed, or Tavily could not parse it | Try running the agent again. If it still fails, verify the server from the MCP Servers page.             |
| Tavily server shows a red status indicator | Server verification failed                           | Open the MCP Servers page, select **Verify** from the Tavily actions menu, and review the error message. |
| Agent uses too many steps and stops early  | Complex page or slow response from Tavily            | Edit the agent and increase **Maximum number of steps** to 50 in General Properties.                     |

## Summary

You have built a fully functional data extraction agent. Here's what you did:

1. ✅ Created a Tavily account and obtained an API key.
2. ✅ Connected Tavily to the platform as an MCP Server with secure API key authentication.
3. ✅ Created an agent with clear extraction instructions and structured output formatting.
4. ✅ Connected the agent to the Tavily extract tool.
5. ✅ Ran the agent against a live web page and received a clean product list.

## Next Steps

* **Extract from your own pages** - replace the sandbox URL with any publicly accessible product page from your own website or a competitor's site, and run the agent again.
* **Schedule regular extractions** - edit the agent's triggers and add a **Schedule** trigger (daily or weekly) so the platform runs the extraction automatically and stores the results in Reports.
* **Chain with a report agent** - create a second agent that takes the product list and formats it as a polished report or comparison table. Connect them so the extraction agent passes its output to the report agent automatically.
* **Expand the instructions** - update the agent's instructions to extract additional fields (e.g. product descriptions, image URLs, or availability status) if the target page includes them.
