外观
概述
使用 LangChain 的嵌入(embeddings)和向量数据库,在 PDF 之上构建一个语义搜索引擎。用它来检索与查询相似的段落,然后将检索器接入检索增强生成(RAG)或其他 LLM 工作流。
本教程涵盖:
- 从 PDF 创建
Document对象。 - 生成嵌入(embeddings)。
- 加载并拆分 PDF。
- 在向量数据库中为文本块建立索引,并按相似度查询。
- 将数据库封装为检索器。
本指南还包含一个基于该搜索引擎构建的最小 RAG 实现。
相关概念
本教程聚焦于文本检索,涵盖以下概念:
环境准备
安装依赖
本教程使用 pypdf 包来读取 PDF:
bash
pip install pypdfbash
conda install pypdf -c conda-forgebash
uv add pypdf本教程使用 pdf-parse 包来读取 PDF:
bash
npm i pdf-parsebash
yarn add pdf-parsebash
pnpm add pdf-parse更多细节,请参阅安装指南。
配置 LangSmith
你使用 LangChain 构建的许多应用程序都会包含多个步骤,并对 LLM 调用多次。 随着这些应用程序变得越来越复杂,能够检视你的链或智能体内部到底发生了什么就变得至关重要。 最好的方式是通过 LangSmith 来实现。
在上面的链接处注册后,请务必设置你的环境变量,以开始记录追踪(trace):
bash
export LANGSMITH_TRACING="true"
export LANGSMITH_API_KEY="..."在 notebook 中,你可以这样设置:
python
import getpass
import os
os.environ["LANGSMITH_TRACING"] = "true"
os.environ["LANGSMITH_API_KEY"] = getpass.getpass()创建文档
LangChain 为一段文本及其关联的元数据实现了 Document 抽象。它有三个属性:
page_content:表示内容的字符串。metadata:包含任意元数据的字典。id:(可选)文档的字符串标识符。pageContent:表示内容的字符串。metadata:包含任意元数据的字典。id:(可选)文档的字符串标识符。
metadata 可以捕获文档的来源、它与其他文档的关系以及其他信息。单个 Document 通常表示更大文档中的一个分块(chunk)。
以下代码创建示例文档:
python
from langchain_core.documents import Document
documents = [
Document(
page_content="Dogs are great companions, known for their loyalty and friendliness.",
metadata={"source": "mammal-pets-doc"},
),
Document(
page_content="Cats are independent pets that often enjoy their own space.",
metadata={"source": "mammal-pets-doc"},
),
]typescript
import { Document } from "@langchain/core/documents";
const documents = [
new Document({
pageContent:
"Dogs are great companions, known for their loyalty and friendliness.",
metadata: { source: "mammal-pets-doc" },
}),
new Document({
pageContent: "Cats are independent pets that often enjoy their own space.",
metadata: { source: "mammal-pets-doc" },
}),
];生成嵌入(embeddings)
向量搜索会存储与文本关联的数值向量。将查询也嵌入为相同维度的向量,然后使用相似度度量(例如余弦相似度)来查找相关文本。
LangChain 支持来自许多提供商的嵌入。选择一个模型来指定如何将文本转换为数值向量:
OpenAI
bash
pip install -U "langchain-openai"python
import getpass
import os
if not os.environ.get("OPENAI_API_KEY"):
os.environ["OPENAI_API_KEY"] = getpass.getpass("Enter API key for OpenAI: ")
from langchain_openai import OpenAIEmbeddings
embeddings = OpenAIEmbeddings(model="text-embedding-3-large")Azure
bash
pip install -U "langchain-openai"python
import getpass
import os
if not os.environ.get("AZURE_OPENAI_API_KEY"):
os.environ["AZURE_OPENAI_API_KEY"] = getpass.getpass("Enter API key for Azure: ")
from langchain_openai import AzureOpenAIEmbeddings
embeddings = AzureOpenAIEmbeddings(
azure_endpoint=os.environ["AZURE_OPENAI_ENDPOINT"],
azure_deployment=os.environ["AZURE_OPENAI_DEPLOYMENT_NAME"],
openai_api_version=os.environ["AZURE_OPENAI_API_VERSION"],
)Google Gemini
bash
pip install -qU langchain-google-genaipython
import getpass
import os
if not os.environ.get("GOOGLE_API_KEY"):
os.environ["GOOGLE_API_KEY"] = getpass.getpass("Enter API key for Google Gemini: ")
from langchain_google_genai import GoogleGenerativeAIEmbeddings
embeddings = GoogleGenerativeAIEmbeddings(model="models/gemini-embedding-001")Google Vertex
bash
pip install -qU langchain-google-vertexaipython
from langchain_google_vertexai import VertexAIEmbeddings
embeddings = VertexAIEmbeddings(model="text-embedding-005")AWS
bash
pip install -qU langchain-awspython
from langchain_aws import BedrockEmbeddings
embeddings = BedrockEmbeddings(model_id="amazon.titan-embed-text-v2:0")HuggingFace
bash
pip install -qU langchain-huggingfacepython
from langchain_huggingface import HuggingFaceEmbeddings
embeddings = HuggingFaceEmbeddings(
model_name="sentence-transformers/all-mpnet-base-v2",
encode_kwargs={"normalize_embeddings": True},
)Ollama
bash
pip install -qU langchain-ollamapython
from langchain_ollama import OllamaEmbeddings
embeddings = OllamaEmbeddings(model="llama3")Cohere
bash
pip install -qU langchain-coherepython
import getpass
import os
if not os.environ.get("COHERE_API_KEY"):
os.environ["COHERE_API_KEY"] = getpass.getpass("Enter API key for Cohere: ")
from langchain_cohere import CohereEmbeddings
embeddings = CohereEmbeddings(model="embed-english-v3.0")MistralAI
bash
pip install -qU langchain-mistralaipython
import getpass
import os
if not os.environ.get("MISTRALAI_API_KEY"):
os.environ["MISTRALAI_API_KEY"] = getpass.getpass("Enter API key for MistralAI: ")
from langchain_mistralai import MistralAIEmbeddings
embeddings = MistralAIEmbeddings(model="mistral-embed")Nomic
bash
pip install -qU langchain-nomicpython
import getpass
import os
if not os.environ.get("NOMIC_API_KEY"):
os.environ["NOMIC_API_KEY"] = getpass.getpass("Enter API key for Nomic: ")
from langchain_nomic import NomicEmbeddings
embeddings = NomicEmbeddings(model="nomic-embed-text-v1.5")NVIDIA
bash
pip install -qU langchain-nvidia-ai-endpointspython
import getpass
import os
if not os.environ.get("NVIDIA_API_KEY"):
os.environ["NVIDIA_API_KEY"] = getpass.getpass("Enter API key for NVIDIA: ")
from langchain_nvidia_ai_endpoints import NVIDIAEmbeddings
embeddings = NVIDIAEmbeddings(model="NV-Embed-QA")Voyage AI
bash
pip install -qU langchain-voyageaipython
import getpass
import os
if not os.environ.get("VOYAGE_API_KEY"):
os.environ["VOYAGE_API_KEY"] = getpass.getpass("Enter API key for Voyage AI: ")
from langchain-voyageai import VoyageAIEmbeddings
embeddings = VoyageAIEmbeddings(model="voyage-3")IBM watsonx
bash
pip install -qU langchain-ibmpython
import getpass
import os
if not os.environ.get("WATSONX_APIKEY"):
os.environ["WATSONX_APIKEY"] = getpass.getpass("Enter API key for IBM watsonx: ")
from langchain_ibm import WatsonxEmbeddings
embeddings = WatsonxEmbeddings(
model_id="ibm/slate-125m-english-rtrvr",
url="https://us-south.ml.cloud.ibm.com",
project_id="<WATSONX PROJECT_ID>",
)Fake
bash
pip install -qU langchain-corepython
from langchain_core.embeddings import DeterministicFakeEmbedding
embeddings = DeterministicFakeEmbedding(size=4096)Isaacus
bash
pip install -qU langchain-isaacuspython
import getpass
import os
if not os.environ.get("ISAACUS_API_KEY"):
os.environ["ISAACUS_API_KEY"] = getpass.getpass("Enter API key for Isaacus: ")
from langchain_isaacus import IsaacusEmbeddings
embeddings = IsaacusEmbeddings(model="kanon-2-embedder")python
vector_1 = embeddings.embed_query(documents[0].page_content)
vector_2 = embeddings.embed_query(documents[1].page_content)
assert len(vector_1) == len(vector_2)
print(f"Generated vectors of length {len(vector_1)}\n")
print(vector_1[:10])OpenAI
bash
npm i @langchain/openaibash
yarn add @langchain/openaibash
pnpm add @langchain/openaitypescript
import { OpenAIEmbeddings } from "@langchain/openai";
const embeddings = new OpenAIEmbeddings({
model: "text-embedding-3-large"
});Azure
bash
npm i @langchain/openaibash
yarn add @langchain/openaibash
pnpm add @langchain/openaibash
AZURE_OPENAI_API_INSTANCE_NAME=<YOUR_INSTANCE_NAME>
AZURE_OPENAI_API_KEY=<YOUR_KEY>
AZURE_OPENAI_API_VERSION="2024-02-01"typescript
import { AzureOpenAIEmbeddings } from "@langchain/openai";
const embeddings = new AzureOpenAIEmbeddings({
azureOpenAIApiEmbeddingsDeploymentName: "text-embedding-ada-002"
});AWS
bash
npm i @langchain/awsbash
yarn add @langchain/awsbash
pnpm add @langchain/awsbash
BEDROCK_AWS_REGION=your-regiontypescript
import { BedrockEmbeddings } from "@langchain/aws";
const embeddings = new BedrockEmbeddings({
model: "amazon.titan-embed-text-v1"
});VertexAI
bash
npm i @langchain/google-vertexaibash
yarn add @langchain/google-vertexaibash
pnpm add @langchain/google-vertexaibash
GOOGLE_APPLICATION_CREDENTIALS=credentials.jsontypescript
import { VertexAIEmbeddings } from "@langchain/google-vertexai";
const embeddings = new VertexAIEmbeddings({
model: "gemini-embedding-001"
});MistralAI
bash
npm i @langchain/mistralaibash
yarn add @langchain/mistralaibash
pnpm add @langchain/mistralaibash
MISTRAL_API_KEY=your-api-keytypescript
import { MistralAIEmbeddings } from "@langchain/mistralai";
const embeddings = new MistralAIEmbeddings({
model: "mistral-embed"
});Cohere
bash
npm i @langchain/coherebash
yarn add @langchain/coherebash
pnpm add @langchain/coherebash
COHERE_API_KEY=your-api-keytypescript
import { CohereEmbeddings } from "@langchain/cohere";
const embeddings = new CohereEmbeddings({
model: "embed-english-v3.0"
});typescript
const vector1 = await embeddings.embedQuery(documents[0].pageContent);
const vector2 = await embeddings.embedQuery(documents[1].pageContent);
assert vector1.length === vector2.length;
console.log(`Generated vectors of length ${vector1.length}\n`);
console.log(vector1.slice(0, 10));txt
Generated vectors of length 1536
[-0.008586574345827103, -0.03341241180896759, -0.008936782367527485, -0.0036674530711025, 0.010564599186182022, 0.009598285891115665, -0.028587326407432556, -0.015824200585484505, 0.0030416189692914486, -0.012899317778646946]接下来,将嵌入存储到支持高效相似度搜索的向量数据库中。
选择向量数据库
LangChain 的 VectorStore 对象会将文本和 Document 对象添加到数据库中,并使用相似度度量进行查询。它们通常使用能将文本转换为数值向量的嵌入模型进行初始化。
LangChain 包含与众多向量数据库技术的集成。其中一些是托管的,需要凭据;一些运行在独立的基础设施中(本地或第三方);另一些则用于轻量级负载,运行在内存中。选择一个向量数据库:
In-memory
bash
pip install -U "langchain-core"python
from langchain_core.vectorstores import InMemoryVectorStore
vector_store = InMemoryVectorStore(embeddings)Amazon OpenSearch
bash
pip install -qU boto3python
from opensearchpy import RequestsHttpConnection
service = "es" # must set the service as 'es'
region = "us-east-2"
credentials = boto3.Session(
aws_access_key_id="xxxxxx", aws_secret_access_key="xxxxx"
).get_credentials()
awsauth = AWS4Auth("xxxxx", "xxxxxx", region, service, session_token=credentials.token)
vector_store = OpenSearchVectorSearch.from_documents(
docs,
embeddings,
opensearch_url="host url",
http_auth=awsauth,
timeout=300,
use_ssl=True,
verify_certs=True,
connection_class=RequestsHttpConnection,
index_name="test-index",
)AstraDB
bash
pip install -U "langchain-astradb"python
from langchain_astradb import AstraDBVectorStore
vector_store = AstraDBVectorStore(
embedding=embeddings,
api_endpoint=ASTRA_DB_API_ENDPOINT,
collection_name="astra_vector_langchain",
token=ASTRA_DB_APPLICATION_TOKEN,
namespace=ASTRA_DB_NAMESPACE,
)Chroma
bash
pip install -qU langchain-chromapython
from langchain_chroma import Chroma
vector_store = Chroma(
collection_name="example_collection",
embedding_function=embeddings,
persist_directory="./chroma_langchain_db", # Where to save data locally, remove if not necessary
)Milvus
bash
pip install -qU langchain-milvuspython
from langchain_milvus import Milvus
URI = "./milvus_example.db"
vector_store = Milvus(
embedding_function=embeddings,
connection_args={"uri": URI},
index_params={"index_type": "FLAT", "metric_type": "L2"},
)MongoDB
bash
pip install -qU langchain-mongodbpython
from langchain_mongodb import MongoDBAtlasVectorSearch
vector_store = MongoDBAtlasVectorSearch(
embedding=embeddings,
collection=MONGODB_COLLECTION,
index_name=ATLAS_VECTOR_SEARCH_INDEX_NAME,
relevance_score_fn="cosine",
)PGVector
bash
pip install -qU langchain-postgrespython
from langchain_postgres import PGVector
vector_store = PGVector(
embeddings=embeddings,
collection_name="my_docs",
connection="postgresql+psycopg://...",
)PGVectorStore
bash
pip install -qU langchain-postgrespython
from langchain_postgres import PGEngine, PGVectorStore
pg_engine = PGEngine.from_connection_string(
url="postgresql+psycopg://..."
)
vector_store = PGVectorStore.create_sync(
engine=pg_engine,
table_name='test_table',
embedding_service=embeddings
)Pinecone
bash
pip install -qU langchain-pineconepython
from langchain_pinecone import PineconeVectorStore
from pinecone import Pinecone
pc = Pinecone(api_key=...)
index = pc.Index(index_name)
vector_store = PineconeVectorStore(embedding=embeddings, index=index)Qdrant
bash
pip install -qU langchain-qdrantpython
from qdrant_client.models import Distance, VectorParams
from langchain_qdrant import QdrantVectorStore
from qdrant_client import QdrantClient
client = QdrantClient(":memory:")
vector_size = len(embeddings.embed_query("sample text"))
if not client.collection_exists("test"):
client.create_collection(
collection_name="test",
vectors_config=VectorParams(size=vector_size, distance=Distance.COSINE)
)
vector_store = QdrantVectorStore(
client=client,
collection_name="test",
embedding=embeddings,
)Memory
bash
npm i @langchain/classicbash
yarn add @langchain/classicbash
pnpm add @langchain/classictypescript
import { MemoryVectorStore } from "@langchain/classic/vectorstores/memory";
const vectorStore = new MemoryVectorStore(embeddings);MongoDB
bash
npm i @langchain/mongodbbash
yarn add @langchain/mongodbbash
pnpm add @langchain/mongodbtypescript
import { MongoDBAtlasVectorSearch } from "@langchain/mongodb"
import { MongoClient } from "mongodb";
const client = new MongoClient(process.env.MONGODB_ATLAS_URI || "");
const collection = client
.db(process.env.MONGODB_ATLAS_DB_NAME)
.collection(process.env.MONGODB_ATLAS_COLLECTION_NAME);
const vectorStore = new MongoDBAtlasVectorSearch(embeddings, {
collection: collection,
indexName: "vector_index",
textKey: "text",
embeddingKey: "embedding",
});Pinecone
bash
npm i @langchain/pineconebash
yarn add @langchain/pineconebash
pnpm add @langchain/pineconetypescript
import { PineconeStore } from "@langchain/pinecone";
import { Pinecone as PineconeClient } from "@pinecone-database/pinecone";
const pinecone = new PineconeClient({
apiKey: process.env.PINECONE_API_KEY,
});
const pineconeIndex = pinecone.Index("your-index-name");
const vectorStore = new PineconeStore(embeddings, {
pineconeIndex,
maxConcurrency: 5,
});Qdrant
bash
npm i @langchain/qdrantbash
yarn add @langchain/qdrantbash
pnpm add @langchain/qdranttypescript
import { QdrantVectorStore } from "@langchain/qdrant";
const vectorStore = await QdrantVectorStore.fromExistingCollection(embeddings, {
url: process.env.QDRANT_URL,
collectionName: "langchainjs-testing",
});Redis
bash
npm i @langchain/redisbash
yarn add @langchain/redisbash
pnpm add @langchain/redistypescript
import { RedisVectorStore } from "@langchain/redis";
const vectorStore = new RedisVectorStore(embeddings, {
redisClient: client,
indexName: "langchainjs-testing",
});加载并拆分 PDF
从 PDF 加载内容,然后在建立索引之前将其拆分成更小的分块。本示例使用一份 2023 年 Nike 10-K 申报文件样本。
python
import pypdf
from langchain_core.documents import Document
# 以下是一个用于演示的最小辅助函数。
def load_pdf_pages(file_path: str) -> list[Document]:
reader = pypdf.PdfReader(file_path)
return [
Document(
page_content=page.extract_text() or "",
metadata={"source": file_path, "page": i},
)
for i, page in enumerate(reader.pages)
]
file_path = "../example_data/nke-10k-2023.pdf"
docs = load_pdf_pages(file_path)
print(len(docs))typescript
import { readFileSync } from "node:fs";
import { Document } from "@langchain/core/documents";
import { PDFParse } from "pdf-parse";
// 以下是一个用于演示的最小辅助函数。
async function loadPdfPages(filePath: string): Promise<Document[]> {
const parser = new PDFParse({
data: new Uint8Array(readFileSync(filePath)),
});
try {
const { pages } = await parser.getText();
return pages.map(
(page) =>
new Document({
pageContent: page.text,
metadata: { source: filePath, page: page.num - 1 },
})
);
} finally {
await parser.destroy();
}
}
const filePath = "../../data/nke-10k-2023.pdf";
const docs = await loadPdfPages(filePath);
console.log(docs.length);txt
107对于检索来说,一页往往过于粗糙。应进一步拆分页面,以免相关的段落被周围的文本稀释。RecursiveCharacterTextSplitter 会递归地在常见的分隔符(如换行符)处拆分,直到每个分块达到目标大小。对于一般的文本用例,这是推荐使用的文本拆分器。
设置 add_start_index=True,这样每次拆分都会保留一个 start_index 元数据字段,用于记录其在原始文档中的字符偏移量。
python
from langchain_text_splitters import RecursiveCharacterTextSplitter
text_splitter = RecursiveCharacterTextSplitter(
chunk_size=1000, chunk_overlap=200, add_start_index=True
)
all_splits = text_splitter.split_documents(docs)
print(len(all_splits))typescript
import { RecursiveCharacterTextSplitter } from "@langchain/textsplitters";
const textSplitter = new RecursiveCharacterTextSplitter({
chunkSize: 1000,
chunkOverlap: 200,
});
const allSplits = await textSplitter.splitDocuments(docs);
console.log(allSplits.length);txt
516为文档建立索引
将分块索引到向量数据库中:
python
ids = vector_store.add_documents(documents=all_splits)typescript
await vectorStore.addDocuments(allSplits);大多数向量数据库集成还支持连接现有的数据库(例如通过客户端或索引名称)。详情请参阅特定集成的文档。
查询向量数据库
在将文档添加到 VectorStore 之后,你就可以查询它了:
- 同步和异步
- 按字符串查询和按向量查询
- 带或不带相似度分数
- 按相似度和最大边际相关性(maximum marginal relevance)查询(以平衡相似度与多样性)
在将文档添加到 VectorStore 之后,你就可以查询它了:
- 同步和异步
- 按字符串查询和按向量查询
- 带或不带相似度分数
- 按相似度和最大边际相关性(maximum marginal relevance)查询(以平衡相似度与多样性)
这些方法通常返回一组 Document 对象。
按字符串搜索
嵌入会将文本映射为稠密向量,使含义相近的内容在几何上彼此靠近。这意味着你可以通过传入一个自然语言问题来检索相关段落:
python
results = vector_store.similarity_search(
"How many distribution centers does Nike have in the US?"
)
print(results[0])python
page_content='direct to consumer operations sell products through the following number of retail stores in the United States:
U.S. RETAIL STORES NUMBER
NIKE Brand factory stores 213
NIKE Brand in-line stores (including employee-only stores) 74
Converse stores (including factory stores) 82
TOTAL 369
In the United States, NIKE has eight significant distribution centers. Refer to Item 2. Properties for further information.
2023 FORM 10-K 2' metadata={'page': 4, 'source': '../example_data/nke-10k-2023.pdf', 'start_index': 3125}typescript
const results1 = await vectorStore.similaritySearch(
"When was Nike incorporated?"
);
console.log(results1[0]);javascript
Document {
pageContent: 'direct to consumer operations sell products...',
metadata: {'page': 4, 'source': '../example_data/nke-10k-2023.pdf', 'start_index': 3125}
}异步查询:
python
results = await vector_store.asimilarity_search("When was Nike incorporated?")
print(results[0])python
page_content='Table of Contents
PART I
ITEM 1. BUSINESS
GENERAL
NIKE, Inc. was incorporated in 1967 under the laws of the State of Oregon. As used in this Annual Report on Form 10-K (this "Annual Report"), the terms "we," "us," "our,"
"NIKE" and the "Company" refer to NIKE, Inc. and its predecessors, subsidiaries and affiliates, collectively, unless the context indicates otherwise.
Our principal business activity is the design, development and worldwide marketing and selling of athletic footwear, apparel, equipment, accessories and services. NIKE is
the largest seller of athletic footwear and apparel in the world. We sell our products through NIKE Direct operations, which are comprised of both NIKE-owned retail stores
and sales through our digital platforms (also referred to as "NIKE Brand Digital"), to retail accounts and to a mix of independent distributors, licensees and sales' metadata={'page': 3, 'source': '../example_data/nke-10k-2023.pdf', 'start_index': 0}返回分数
你可以随文档一起返回相似度分数。分数的含义因提供商而异。在本例中,分数是一种距离度量,与相似度成反比:
python
# 请注意,各提供商实现不同的分数;这里的分数
# 是一种与相似度成反比的距离度量。
results = vector_store.similarity_search_with_score("What was Nike's revenue in 2023?")
doc, score = results[0]
print(f"Score: {score}\n")
print(doc)python
Score: 0.23699893057346344
page_content='Table of Contents
FISCAL 2023 NIKE BRAND REVENUE HIGHLIGHTS
The following tables present NIKE Brand revenues disaggregated by reportable operating segment, distribution channel and major product line:
FISCAL 2023 COMPARED TO FISCAL 2022
•NIKE, Inc. Revenues were $51.2 billion in fiscal 2023, which increased 10% and 16% compared to fiscal 2022 on a reported and currency-neutral basis, respectively.
The increase was due to higher revenues in North America, Europe, Middle East & Africa ("EMEA"), APLA and Greater China, which contributed approximately 7, 6,
2 and 1 percentage points to NIKE, Inc. Revenues, respectively.
•NIKE Brand revenues, which represented over 90% of NIKE, Inc. Revenues, increased 10% and 16% on a reported and currency-neutral basis, respectively. This
increase was primarily due to higher revenues in Men's, the Jordan Brand, Women's and Kids' which grew 17%, 35%,11% and 10%, respectively, on a wholesale
equivalent basis.' metadata={'page': 35, 'source': '../example_data/nke-10k-2023.pdf', 'start_index': 0}typescript
const results2 = await vectorStore.similaritySearchWithScore(
"What was Nike's revenue in 2023?"
);
console.log(results2[0]);javascript
Score: 0.23699893057346344
Document {
pageContent: 'Table of Contents...',
metadata: {'page': 35, 'source': '../example_data/nke-10k-2023.pdf', 'start_index': 0}
}按向量搜索
自己嵌入查询,然后用得到的向量进行搜索:
python
embedding = embeddings.embed_query("How were Nike's margins impacted in 2023?")
results = vector_store.similarity_search_by_vector(embedding)
print(results[0])python
page_content='Table of Contents
GROSS MARGIN
FISCAL 2023 COMPARED TO FISCAL 2022
For fiscal 2023, our consolidated gross profit increased 4% to $22,292 million compared to $21,479 million for fiscal 2022. Gross margin decreased 250 basis points to
43.5% for fiscal 2023 compared to 46.0% for fiscal 2022 due to the following:
*Wholesale equivalent
The decrease in gross margin for fiscal 2023 was primarily due to:
•Higher NIKE Brand product costs, on a wholesale equivalent basis, primarily due to higher input costs and elevated inbound freight and logistics costs as well as
product mix;
•Lower margin in our NIKE Direct business, driven by higher promotional activity to liquidate inventory in the current period compared to lower promotional activity in
the prior period resulting from lower available inventory supply;
•Unfavorable changes in net foreign currency exchange rates, including hedges; and
•Lower off-price margin, on a wholesale equivalent basis.
This was partially offset by:' metadata={'page': 36, 'source': '../example_data/nke-10k-2023.pdf', 'start_index': 0}typescript
const embedding = await embeddings.embedQuery(
"How were Nike's margins impacted in 2023?"
);
const results3 = await vectorStore.similaritySearchVectorWithScore(
embedding,
1
);
console.log(results3[0]);javascript
Document {
pageContent: 'FISCAL 2023 COMPARED TO FISCAL 2022...',
metadata: {
'page': 36,
'source': '../example_data/nke-10k-2023.pdf',
'start_index': 0
}
}了解更多:
- API Reference
- 集成相关文档
使用检索器
LangChain 的 VectorStore 对象并不继承 Runnable。而检索器(Retrievers)是 runnable,因此它们支持标准方法,例如同步和异步的 invoke 与 batch。
你也可以从向量数据库构建检索器,检索器还可以封装非向量来源(例如外部 API)。
在本例中,通过封装 similarity_search 来创建一个简单的检索器,无需继承 Retriever:
python
from typing import List
from langchain_core.documents import Document
from langchain_core.runnables import chain
@chain
def retriever(query: str) -> List[Document]:
return vector_store.similarity_search(query, k=1)
retriever.batch(
[
"How many distribution centers does Nike have in the US?",
"When was Nike incorporated?",
],
)txt
[[Document(metadata={'page': 4, 'source': '../example_data/nke-10k-2023.pdf', 'start_index': 3125}, page_content='direct to consumer operations sell products through the following number of retail stores in the United States:\nU.S. RETAIL STORES NUMBER\nNIKE Brand factory stores 213 \nNIKE Brand in-line stores (including employee-only stores) 74 \nConverse stores (including factory stores) 82 \nTOTAL 369 \nIn the United States, NIKE has eight significant distribution centers. Refer to Item 2. Properties for further information.\n2023 FORM 10-K 2')],
[Document(metadata={'page': 3, 'source': '../example_data/nke-10k-2023.pdf', 'start_index': 0}, page_content='Table of Contents\nPART I\nITEM 1. BUSINESS\nGENERAL\nNIKE, Inc. was incorporated in 1967 under the laws of the State of Oregon. As used in this Annual Report on Form 10-K (this "Annual Report"), the terms "we," "us," "our,"\n"NIKE" and the "Company" refer to NIKE, Inc. and its predecessors, subsidiaries and affiliates, collectively, unless the context indicates otherwise.\nOur principal business activity is the design, development and worldwide marketing and selling of athletic footwear, apparel, equipment, accessories and services. NIKE is\nthe largest seller of athletic footwear and apparel in the world. We sell our products through NIKE Direct operations, which are comprised of both NIKE-owned retail stores\nand sales through our digital platforms (also referred to as "NIKE Brand Digital"), to retail accounts and to a mix of independent distributors, licensees and sales')]]向量数据库实现了 as_retriever 方法,返回一个 VectorStoreRetriever。这些检索器提供 search_type 和 search_kwargs,用于选择并参数化底层存储的方法。用下面的方式复现上面的示例:
python
retriever = vector_store.as_retriever(
search_type="similarity",
search_kwargs={"k": 1},
)
retriever.batch(
[
"How many distribution centers does Nike have in the US?",
"When was Nike incorporated?",
],
)txt
[[Document(metadata={'page': 4, 'source': '../example_data/nke-10k-2023.pdf', 'start_index': 3125}, page_content='direct to consumer operations sell products through the following number of retail stores in the United States:\nU.S. RETAIL STORES NUMBER\nNIKE Brand factory stores 213 \nNIKE Brand in-line stores (including employee-only stores) 74 \nConverse stores (including factory stores) 82 \nTOTAL 369 \nIn the United States, NIKE has eight significant distribution centers. Refer to Item 2. Properties for further information.\n2023 FORM 10-K 2')],
[Document(metadata={'page': 3, 'source': '../example_data/nke-10k-2023.pdf', 'start_index': 0}, page_content='Table of Contents\nPART I\nITEM 1. BUSINESS\nGENERAL\nNIKE, Inc. was incorporated in 1967 under the laws of the State of Oregon. As used in this Annual Report on Form 10-K (this "Annual Report"), the terms "we," "us," "our,"\n"NIKE" and the "Company" refer to NIKE, Inc. and its predecessors, subsidiaries and affiliates, collectively, unless the context indicates otherwise.\nOur principal business activity is the design, development and worldwide marketing and selling of athletic footwear, apparel, equipment, accessories and services. NIKE is\nthe largest seller of athletic footwear and apparel in the world. We sell our products through NIKE Direct operations, which are comprised of both NIKE-owned retail stores\nand sales through our digital platforms (also referred to as "NIKE Brand Digital"), to retail accounts and to a mix of independent distributors, licensees and sales')]]VectorStoreRetriever 支持 "similarity"(默认)、"mmr"(最大边际相关性)和 "similarity_score_threshold" 等搜索类型。使用最后一种选项可以按相似度分数过滤文档。
typescript
const retriever = vectorStore.asRetriever({
searchType: "mmr",
searchKwargs: {
fetchK: 1,
},
});
await retriever.batch([
"When was Nike incorporated?",
"What was Nike's revenue in 2023?",
]);javascript
[
[Document {
metadata: {'page': 4, 'source': '../example_data/nke-10k-2023.pdf', 'start_index': 3125},
pageContent: 'direct to consumer operations sell products...',
}],
[Document {
metadata: {'page': 3, 'source': '../example_data/nke-10k-2023.pdf', 'start_index': 0},
pageContent: 'Table of Contents...',
}],
]你可以在更复杂的应用中使用检索器,例如检索增强生成(RAG),它会在提示词中将一个问题与检索到的上下文结合起来交给 LLM。要了解更多关于构建此类应用的内容,请查看 RAG 教程。
后续步骤
你现在已经了解了如何在 PDF 文档之上构建语义搜索引擎。
更多信息请参阅:
关于 RAG 的更多内容: