diff --git a/img/vector/create_index_with_model.png b/img/vector/create_index_with_model.png
index 341ecd88f..070e8592b 100644
Binary files a/img/vector/create_index_with_model.png and b/img/vector/create_index_with_model.png differ
diff --git a/llms-full.txt b/llms-full.txt
index 691a85e4d..1efbde2e1 100644
--- a/llms-full.txt
+++ b/llms-full.txt
@@ -99184,9 +99184,9 @@ curl $UPSTASH_VECTOR_REST_URL/info \
"similarityFunction": "COSINE",
"indexType": "HYBRID",
"denseIndex": {
- "dimension": 1024,
+ "dimension": 1536,
"similarityFunction": "COSINE",
- "embeddingModel": "BGE_M3"
+ "embeddingModel": "TEXT_EMBEDDING_3_SMALL"
},
"sparseIndex": {
"embeddingModel": "BM25"
@@ -100649,56 +100649,38 @@ you can now upsert and query raw string data when using your database instead of
converting your text to a vector first. The vectorization is done automatically
by your selected model.
-## Upstash Embedding Models - Video Guide
-
-Let's look at how Upstash embeddings work, how the models we offer compare, and
-which model is best for your use case.
-
-
-
## Models
-Upstash Vector comes with a variety of embedding models that score well in the
-[MTEB](https://huggingface.co/spaces/mteb/leaderboard) leaderboard, a benchmark
-for measuring the performance of embedding models. They support use cases such
-as classification, clustering, or retrieval.
+Upstash Vector hosts the following embedding model for dense and hybrid indexes:
-You can choose the following general purpose models for dense and hybrid indexes:
+| Name | Dimension | Sequence Length | MTEB |
+| ----------------------------------------------------------------------------------------- | --------- | --------------- | ---- |
+| [openai/text-embedding-3-small](https://platform.openai.com/docs/guides/embeddings) | 1536 | 8191 | 62.3 |
-| Name | Dimension | Sequence Length | MTEB |
-| ------------------------------------------------------------------------------------------------------- | --------- | --------------- | ----- |
-| [BAAI/bge-large-en-v1.5](https://huggingface.co/BAAI/bge-large-en-v1.5) | 1024 | 512 | 64.23 |
-| [BAAI/bge-base-en-v1.5](https://huggingface.co/BAAI/bge-base-en-v1.5) | 768 | 512 | 63.55 |
-| [BAAI/bge-small-en-v1.5](https://huggingface.co/BAAI/bge-small-en-v1.5) | 384 | 512 | 62.17 |
-| [BAAI/bge-m3](https://huggingface.co/BAAI/bge-m3) | 1024 | 8192 | * |
+The MTEB score is the average reported by OpenAI in its
+[model announcement](https://openai.com/index/new-embedding-models-and-api-updates/).
- The sequence length is not a hard limit. Models truncate the input
- appropriately when given a raw text data that would result in more tokens than
- the given sequence length. However, we recommend using appropriate models and
- not exceeding their sequence length to have more accurate results.
+ The sequence length is not a hard limit. The model truncates the input
+ appropriately when given raw text that would result in more tokens than
+ the given sequence length. However, we recommend not exceeding the sequence
+ length to have more accurate results.
-
- MTEB score for the `BAAI/bge-m3` is not fully measured.
-
+For sparse and hybrid indexes, the following model can be selected:
-For sparse and hybrid indexes, on the following models can be selected:
+| Name |
+| ------------------------------------------------ |
+| [BM25](https://en.wikipedia.org/wiki/Okapi_BM25) |
-| Name |
-| ------------------------------------------------- |
-| [BAAI/bge-m3](https://huggingface.co/BAAI/bge-m3) |
-| [BM25](https://en.wikipedia.org/wiki/Okapi_BM25) |
+See [Creating Sparse Vectors](/docs/vector/features/sparseindexes#creating-sparse-vectors) for the details of the above model.
-See [Creating Sparse Vectors](/docs/vector/features/sparseindexes#creating-sparse-vectors) for the details of the above models.
+
+ The BGE models (`BAAI/bge-large-en-v1.5`, `BAAI/bge-base-en-v1.5`,
+ `BAAI/bge-small-en-v1.5` and `BAAI/bge-m3`) are no longer available for new
+ indexes. If you need a different model, you can generate the embeddings
+ yourself and create the index with a custom dimension.
+
## Using a Model
@@ -103868,7 +103850,7 @@ Sparse vectors are representations in a high-dimensional space,
where only a small number of dimensions have non-zero values.
For example, for the same text, a dense vector representation with
-the BGE-M3 model would have 1024 non-zero valued dimensions.
+the `text-embedding-3-small` model would have values in all of its 1536 dimensions.
However, the sparse vector representation of the same text would
have less than a hundred non-zero valued dimensions, whereas vector
space potentially has more than 250 thousand dimensions. Also, unlike
@@ -103917,24 +103899,12 @@ that enhance documents and queries with term weighting and expansion.
Upstash gives you full control by allowing you to upsert and query
sparse vectors.
-Also, to make embedding easier for you, Upstash provides some hosted
-models and allows you to upsert and query text data. Behind the scenes,
+Also, to make embedding easier for you, Upstash provides a hosted
+BM25 model and allows you to upsert and query text data. Behind the scenes,
the text data is converted to sparse vectors.
You can create your index with a sparse embedding model to use this feature.
-### BGE-M3 Sparse Vectors
-
-BGE-M3 is a multi-functional, multi-lingual, and multi-granular model
-widely used for dense indexes.
-
-We also provide BGE-M3 as a sparse vector embedder, which outputs
-sparse vectors from `250_002` dimensional space.
-
-These sparse vectors have values where each token is weighted
-according to the input text, which enhances traditional sparse vectors
-with contextuality.
-
### BM25 Sparse Vectors
BM25 is a popular algorithm used in full-text search systems to rank
diff --git a/vector/api/endpoints/info.mdx b/vector/api/endpoints/info.mdx
index 3be606680..aa499152b 100644
--- a/vector/api/endpoints/info.mdx
+++ b/vector/api/endpoints/info.mdx
@@ -94,9 +94,9 @@ curl $UPSTASH_VECTOR_REST_URL/info \
"similarityFunction": "COSINE",
"indexType": "HYBRID",
"denseIndex": {
- "dimension": 1024,
+ "dimension": 1536,
"similarityFunction": "COSINE",
- "embeddingModel": "BGE_M3"
+ "embeddingModel": "TEXT_EMBEDDING_3_SMALL"
},
"sparseIndex": {
"embeddingModel": "BM25"
diff --git a/vector/features/embeddingmodels.mdx b/vector/features/embeddingmodels.mdx
index 93f355227..994f54d50 100644
--- a/vector/features/embeddingmodels.mdx
+++ b/vector/features/embeddingmodels.mdx
@@ -11,56 +11,38 @@ you can now upsert and query raw string data when using your database instead of
converting your text to a vector first. The vectorization is done automatically
by your selected model.
-## Upstash Embedding Models - Video Guide
-
-Let's look at how Upstash embeddings work, how the models we offer compare, and
-which model is best for your use case.
-
-
-
## Models
-Upstash Vector comes with a variety of embedding models that score well in the
-[MTEB](https://huggingface.co/spaces/mteb/leaderboard) leaderboard, a benchmark
-for measuring the performance of embedding models. They support use cases such
-as classification, clustering, or retrieval.
+Upstash Vector hosts the following embedding model for dense and hybrid indexes:
-You can choose the following general purpose models for dense and hybrid indexes:
+| Name | Dimension | Sequence Length | MTEB |
+| ----------------------------------------------------------------------------------------- | --------- | --------------- | ---- |
+| [openai/text-embedding-3-small](https://platform.openai.com/docs/guides/embeddings) | 1536 | 8191 | 62.3 |
-| Name | Dimension | Sequence Length | MTEB |
-| ------------------------------------------------------------------------------------------------------- | --------- | --------------- | ----- |
-| [BAAI/bge-large-en-v1.5](https://huggingface.co/BAAI/bge-large-en-v1.5) | 1024 | 512 | 64.23 |
-| [BAAI/bge-base-en-v1.5](https://huggingface.co/BAAI/bge-base-en-v1.5) | 768 | 512 | 63.55 |
-| [BAAI/bge-small-en-v1.5](https://huggingface.co/BAAI/bge-small-en-v1.5) | 384 | 512 | 62.17 |
-| [BAAI/bge-m3](https://huggingface.co/BAAI/bge-m3) | 1024 | 8192 | * |
+The MTEB score is the average reported by OpenAI in its
+[model announcement](https://openai.com/index/new-embedding-models-and-api-updates/).
- The sequence length is not a hard limit. Models truncate the input
- appropriately when given a raw text data that would result in more tokens than
- the given sequence length. However, we recommend using appropriate models and
- not exceeding their sequence length to have more accurate results.
+ The sequence length is not a hard limit. The model truncates the input
+ appropriately when given raw text that would result in more tokens than
+ the given sequence length. However, we recommend not exceeding the sequence
+ length to have more accurate results.
-
- MTEB score for the `BAAI/bge-m3` is not fully measured.
-
+For sparse and hybrid indexes, the following model can be selected:
-For sparse and hybrid indexes, on the following models can be selected:
+| Name |
+| ------------------------------------------------ |
+| [BM25](https://en.wikipedia.org/wiki/Okapi_BM25) |
-| Name |
-| ------------------------------------------------- |
-| [BAAI/bge-m3](https://huggingface.co/BAAI/bge-m3) |
-| [BM25](https://en.wikipedia.org/wiki/Okapi_BM25) |
+See [Creating Sparse Vectors](/vector/features/sparseindexes#creating-sparse-vectors) for the details of the above model.
-See [Creating Sparse Vectors](/vector/features/sparseindexes#creating-sparse-vectors) for the details of the above models.
+
+ The BGE models (`BAAI/bge-large-en-v1.5`, `BAAI/bge-base-en-v1.5`,
+ `BAAI/bge-small-en-v1.5` and `BAAI/bge-m3`) are no longer available for new
+ indexes. If you need a different model, you can generate the embeddings
+ yourself and create the index with a custom dimension.
+
## Using a Model
diff --git a/vector/features/sparseindexes.mdx b/vector/features/sparseindexes.mdx
index aa8ddfefd..5ae1d4a35 100644
--- a/vector/features/sparseindexes.mdx
+++ b/vector/features/sparseindexes.mdx
@@ -6,7 +6,7 @@ Sparse vectors are representations in a high-dimensional space,
where only a small number of dimensions have non-zero values.
For example, for the same text, a dense vector representation with
-the BGE-M3 model would have 1024 non-zero valued dimensions.
+the `text-embedding-3-small` model would have values in all of its 1536 dimensions.
However, the sparse vector representation of the same text would
have less than a hundred non-zero valued dimensions, whereas vector
space potentially has more than 250 thousand dimensions. Also, unlike
@@ -55,24 +55,12 @@ that enhance documents and queries with term weighting and expansion.
Upstash gives you full control by allowing you to upsert and query
sparse vectors.
-Also, to make embedding easier for you, Upstash provides some hosted
-models and allows you to upsert and query text data. Behind the scenes,
+Also, to make embedding easier for you, Upstash provides a hosted
+BM25 model and allows you to upsert and query text data. Behind the scenes,
the text data is converted to sparse vectors.
You can create your index with a sparse embedding model to use this feature.
-### BGE-M3 Sparse Vectors
-
-BGE-M3 is a multi-functional, multi-lingual, and multi-granular model
-widely used for dense indexes.
-
-We also provide BGE-M3 as a sparse vector embedder, which outputs
-sparse vectors from `250_002` dimensional space.
-
-These sparse vectors have values where each token is weighted
-according to the input text, which enhances traditional sparse vectors
-with contextuality.
-
### BM25 Sparse Vectors
BM25 is a popular algorithm used in full-text search systems to rank