Skip to main content
Version: August 2026 - Dify 1.16.1

โš™๏ธ Exercise 2.1: Ingest Knowledge Using the API

๐ŸŽฏ Exercise Overview

Go beyond manual upload! In this exercise, you'll use a ready-to-use API to automatically ingest Jira issues into your Dify knowledge base. Instead of one big text file, each issue becomes its own document, tagged with metadata such as its issue_key. The same approach works for bulk uploads and scheduled pipelines at work.

โœจ This is the foundation for automated knowledge ingestion. Perfect for teams that want to integrate Jira data into their AI assistants!

๐Ÿš€ What You'll Build

  • Connect to the deployed Jira ingestion API
  • Explore available Jira projects using Swagger UI
  • Ingest structured Jira issues into your Dify knowledge base, in a basic and an advanced mode
  • Improve retrieval with Summary Index and hybrid search
  • Compare manual vs. API-based ingestion results
Knowledge Base

๐Ÿ“‹ Step-by-Step Checklistโ€‹

๐ŸŽฏ Exercise Checklist

โฑ๏ธ Estimated Time: 20-25 minutes | ๐ŸŽฏ Goal: API-ingested knowledge base with structured metadata and measured retrieval


๐Ÿ› ๏ธ Step 1: Get Your Dify Configurationโ€‹

๐Ÿ“‹ Configuration Steps

You'll need your Dify instance details to connect the API. Follow these steps:

1.1: Access Your Dify API Settingsโ€‹

  1. In your Dify instance, navigate to Knowledge section

  2. Find and copy your API Server URL

  3. Click the API Key button and create a dedicated, short-lived key for this ingestion

Credential safety: Use this key only for the workshop ingestion, never include it in screenshots or saved command files, and revoke it after the verification steps below.

API Access
Step 1: Access API settings in Knowledge section (click to expand)
Create API Key
Step 2: Create API key for datasets (click to expand)

๐Ÿง  Pro Tip: Keep your API Server URL and API key handy - you'll need them in the next steps!


๐Ÿ” Step 2: Explore Available Projectsโ€‹

Before sending any Dify credentials, check that the service is up: open /health (it should respond OK) and /projects (it should list REST_JiraEcosystem_issues).

๐Ÿ“‹ Exploration Steps

Before ingesting data, let's see what Jira projects are available using the interactive Swagger UI.

2.1: Open the API Documentationโ€‹

๐Ÿ“š Interactive API Documentation

Explore available Jira projects and test API endpoints directly in your browser

๐Ÿ”— Open Swagger UI โ†’

The Swagger UI provides an interactive interface where you can:

  • ๐Ÿ“– Browse all available endpoints
  • ๐Ÿงช Test API calls directly in your browser
  • ๐Ÿ‘€ See request/response examples
  • ๐Ÿ” View available projects interactively
Swagger UI showing available projects
Swagger UI interface (click to expand)

๐Ÿ’ก Tip: The Swagger UI is especially helpful if you're new to APIs. You can click "Try it out" on any endpoint to test it without writing curl commands!


๐Ÿ“ค Step 3: Ingest Jira Issuesโ€‹

๐Ÿ”ง Ingestion Steps

Use the POST /jira/ingest endpoint to ingest specific projects. Leave the dataset_id empty and fill in your Dify configuration.

3.1: Using Command Line (curl)โ€‹

Run this command to ingest a project:

curl -X 'POST' \
'https://dify-jira.testingfantasy.com/jira/ingest?advanced_ingestion=false' \
-H 'accept: application/json' \
-H 'Content-Type: application/json' \
-d '{
"request": {
"project": "REST_JiraEcosystem_issues"
},
"dify_config": {
"dify_base_url": "{your_dify_base_url}/v1",
"dify_api_key": "{your_dify_api_key}",
"dataset_id": ""
}
}'
๐Ÿ“ Replace these values:
  • {your_dify_base_url} โ†’ Your Dify instance base URL (from Step 1)
  • {your_dify_api_key} โ†’ Your Dify API key (from Step 1)
  • REST_JiraEcosystem_issues โ†’ The project you want to ingest (use /projects endpoint to see all available projects)

You can also use the Swagger UI to test the endpoint interactively:

  1. Navigate to the POST /jira/ingest endpoint in Swagger UI
  2. Click "Try it out"
  3. Use this payload:
{
"request": {
"project": "REST_JiraEcosystem_issues"
},
"dify_config": {
"dify_base_url": "{your_dify_base_url}/v1",
"dify_api_key": "{your_dify_api_key}",
"dataset_id": ""
}
}
  1. Click "Execute" to run the ingestion
Swagger UI showing ingest endpoint
Swagger UI ingest endpoint (click to expand)

3.3: Run the Advanced Ingestionโ€‹

Run the same request a second time, now with advanced_ingestion=true. This creates a second knowledge base, Jira_API_Advanced_*, so you can compare the two in Step 7.

curl -X 'POST' \
'https://dify-jira.testingfantasy.com/jira/ingest?advanced_ingestion=true' \
-H 'accept: application/json' \
-H 'Content-Type: application/json' \
-d '{
"request": {
"project": "REST_JiraEcosystem_issues"
},
"dify_config": {
"dify_base_url": "{your_dify_base_url}/v1",
"dify_api_key": "{your_dify_api_key}",
"dataset_id": ""
}
}'

In Swagger UI, set advanced_ingestion to true and execute the same payload.

What's different? Advanced ingestion enriches every issue document before it's indexed:

  • Aliases: the other ways people refer to the issue, such as REST-266, Issue REST-266, and Jira REST-266.
  • Example queries: questions a user might ask about the issue. They give the retriever more wording to match against.

This is called query seeding. It helps when the user's question doesn't use the same words as the Jira issue.


โœ… Step 4: Verify Ingestion in Difyโ€‹

๐Ÿ” Verification Steps

After running the ingestion command, verify that your data was successfully uploaded to Dify.

4.1: Check the API Responseโ€‹

The API will return a response showing the ingestion status:

{
"success": true,
"dify_instance": "{your_dify_base_url}/v1",
"files_processed": 1,
"files_failed": 0,
"results": [
{
"file": "REST_JiraEcosystem_issues",
"result": ["... Dify document and metadata responses ..."],
"issues_ingested": 23
}
],
"errors": []
}

Treat the request as successful only when success is true, files_processed is 1, files_failed is 0, and errors is empty.

4.2: Verify in Dify Knowledge Baseโ€‹

Navigate to your Dify instance and check the Knowledge section. You should see your two new knowledge bases, Jira_API_Basic_* and Jira_API_Advanced_*.

Confirm that:

  • both knowledge bases exist and all 23 documents in each show Available;
  • built-in metadata is enabled and the custom issue_key field is attached to issue documents; and
  • the descriptive smoke query REST-266 Default Jackson configuration unknown properties returns REST-266 at rank 1.

Indexing continues in Dify after the API responds, so check the documents in Dify, not just the API response. Every run creates a new knowledge base: don't repeat a call unless you want another copy.

KB Created API
Knowledge base created via API (click to expand)
Dataset API Config
Dataset configuration (click to expand)

๐ŸŽ‰ Success! Your Knowledge Base is Ready

Your Jira issues have been successfully ingested with structured metadata!

After the knowledge-base and retrieval checks pass, return to Dify's Knowledge API settings and revoke the short-lived key. Do not keep the key in workshop notes, terminal history, screenshots, or source control.


๐Ÿง  Step 5: Add Summary Indexโ€‹

Dify's Summary Index generates a compact LLM summary for each chunk and embeds that summary as another retrieval surface. This can make requirements, risks, and test implications easier to find when the original Jira wording differs from the learner's question.

  1. Open the Jira_API_Basic_* knowledge base and go to Settings.
  2. Keep High Quality indexing and enable Summary Auto-Gen.
  3. Select gpt-35-turbo-16k.
  4. Instruct the model to preserve the Jira key, summary, status, requirements, risks, acceptance criteria, test implications, and constraints without inventing facts.
  5. Save the settings. Summary Auto-Gen applies to newly indexed content; for the 23 existing documents, select all documents and choose Generate summary.
Summary Auto-Gen enabled with the workshop model and Jira-focused instructions

Inspect at least one generated summary before trusting it. For REST-266, confirm that the backward-compatibility constraint remains explicit. Generated test implications are useful retrieval hints, but they are model output, not new Jira requirements.


๐ŸŽ›๏ธ Step 6: Tune and Benchmark Retrievalโ€‹

Configure Hybrid Search with weighted retrieval:

  • semantic weight: 0.3;
  • keyword weight: 0.7;
  • Top K: 10; and
  • score threshold: 0.05.
Hybrid Search weighted toward keywords with a 0.05 score threshold

Run this fixed acceptance suite in Retrieval Testing before tuning further:

QueryExpected evidence
REST-266 Default Jackson configuration unknown propertiesREST-266 at rank 1
Which issue creates a backward-compatibility risk when unknown JSON properties are accepted?REST-266 at rank 1
Which Jira work has important test implications for REST API compatibility?REST-266 in the top 3
Why is the sky blue?No retrieved chunks
Risk-oriented benchmark retrieving REST-266 at rank 1 Unrelated benchmark query returning no chunks after the score threshold

Hybrid search does not replace metadata filtering. Conceptual Jira questions benefit from summaries and semantic retrieval, while exact identifiers such as REST-265 should be extracted and applied as an issue_key filter. Exercise 3 establishes the grounded baseline; Exercise 4 adds the extractor and exact-key retrieval paths.


๐Ÿ“š Step 7: Compare Knowledge Base Approachesโ€‹

๐Ÿ”ฌ Understanding the Differences

You now have three knowledge bases in Dify:

๐Ÿ“„ Manual Ingestion

From Exercise 2: .txt files uploaded by hand. Simple, but limited structure and no metadata.

๐Ÿ”ง Basic API Ingestion

Jira_API_Basic_* from Step 3: one document per issue, with metadata. Useful for simple lookups and basic retrieval.

๐Ÿš€ Advanced API Ingestion

Jira_API_Advanced_* from Step 3.3: adds aliases and example queries. Better at matching different wordings of the same question.

๐Ÿ’ก Key Differences:
  • Manual: Raw text files, limited structure, no metadata
  • Basic API: Structured data with metadata (issue keys, project info)
  • Advanced API: Query seeding + aliases for better retrieval and understanding
๐Ÿงช Test These Queries on All Three Knowledge Bases:
  • ๐Ÿ‘‰ What is this issue about: Jira Issue: REST-259?
  • ๐Ÿ‘‰ Give a test plan for REST-266
  • ๐Ÿ‘‰ Who is the assignee for REST-266?
  • ๐Ÿ‘‰ Summarize the technical documentation for the REST module.

Record rank, source, score, and no-match behavior rather than judging only how plausible a result looks.

Here is What is issue 266 about? in Retrieval Testing on each knowledge base, with the default vector search:

๐Ÿ“„ Manual: the top chunk is an unrelated slice of the .txt file. The words "issue 266" don't match anything meaningful.

Manual knowledge base returning an unrelated chunk from REST_JiraEcosystem_issues.txt with score 0.26

๐Ÿ”ง Basic API: each result is one clean Jira issue, but the top match is REST-432, the wrong one. The number 266 alone carries little meaning for the embedding.

Basic API knowledge base returning REST-432 and REST-259 for a question about issue 266

๐Ÿš€ Advanced API: REST-266 comes first. Its aliases (Issue 266, Jira 266) and example questions give the retriever the words users actually type.

Advanced API knowledge base returning REST-266 first, with its aliases, score 0.34

๐Ÿง  Pro Tip: Open Retrieval Testing in each knowledge base and try the queries above. Compare the answers you get from each knowledge base. Structured data and advanced techniques enable smarter, more accurate answers!


๐Ÿ” What Happens Behind the Scenes?โ€‹

๐Ÿ”ฌ Technical Details

When you ingest Jira issues via the API, here's what happens:

  • Each issue is converted into a descriptive document with summary, description, project, type, etc.
  • Smart chunking ensures readable chunks with overlap for better search coverage
  • Metadata (like issue_key) is attached to each chunk, improving traceability
  • Documents are vectorized and indexed in your Dify knowledge base

๐Ÿ”ง Advanced: Deploy Your Own API (Optional)โ€‹

๐Ÿš€ Deploy Your Own API Instance

If you want to deploy your own version of the API or understand how it works under the hood:

  1. Visit https://github.com/bassagap/dify_jira and create a Codespace on the main branch
  2. Follow the setup instructions to deploy your own instance
  3. This gives you full control over the API and allows customization

Note: This is optional and only needed if you want to customize the API or understand the implementation details.


๐ŸŽฏ Exercise Complete! What's Next?โ€‹

๐ŸŽ‰ Congratulations!

You've successfully ingested Jira issues using the API!

โœ… What You've Accomplished:

  • โœ… Connected to the Jira ingestion API
  • โœ… Explored available projects using Swagger UI
  • โœ… Ingested structured Jira issues with metadata, in basic and advanced mode
  • โœ… Compared different knowledge base approaches
  • โœ… Generated and reviewed Summary Index content
  • โœ… Passed a fixed positive and no-match retrieval benchmark

๐Ÿš€ Ready for the Next Challenge?

Your knowledge bases can find the right Jira issues. In Exercise 3, you'll connect one to a chatbot so it answers questions grounded in those issues.