pgvector and Cloudberry Database version
Environment
- OS: Ubuntu 24.04.3 LTS (Noble Numbat)
- Kernel: Linux node195 6.8.0-117-generic
- Architecture: x86_64
- Database: Apache Cloudberry
- Cluster layout: single-machine Apache Cloudberry cluster
1 coordinator
2 primary segments
observed segment logs from seg0 127.0.1.1:6000 and seg1 127.0.1.1:6001
psql --version:
psql (Apache Cloudberry) 14.4
pg_config --version:
PostgreSQL 14.4
SELECT version();:
PostgreSQL 14.4 (Apache Cloudberry 3.0.0-devel+dev.2227.g017b06b4f86 build dev) on x86_64-pc-linux-gnu, compiled by gcc-11 (Ubuntu 11.5.0-1ubuntu1~24.04) 11.5.0, 64-bit compiled on Feb 2 2026 05:21:20 (with assert checking)
SHOW server_version;:
14.4
pgvector extension version: 0.8.0
~/projects/pgvector$ git remote -v
origin git@github.com:cloudberry-contrib/pgvector.git (fetch)
origin git@github.com:cloudberry-contrib/pgvector.git (push)
What happened
Summary
When creating a pgvector HNSW index on Apache Cloudberry with parallel maintenance workers enabled, the index build enters the parallel worker path and then crashes the server process.
The log shows pgvector starts using parallel workers:
DEBUG: using 2 parallel workers
but the parallel worker later fails with a Cloudberry interconnect assertion:
FailedAssertion("ht->size > 0", File: "udp/ic_udpifc.c", Line: 2013)
The client connection is then terminated because another server process crashed.
### What you think should happen instead
It looks like the HNSW index build entered pgvector's parallel build path, since the log contains `using 2 parallel workers`. The failure then happened inside a parallel worker with a Cloudberry interconnect assertion:
`FailedAssertion("ht->size > 0", File: "udp/ic_udpifc.c", Line: 2013)`
My guess is that pgvector's PostgreSQL-style parallel index build path is not fully compatible with Cloudberry's segment/interconnect execution model in this case. The confusing part is that the log also says the index is being built "serially" before reporting `using 2 parallel workers`, so maybe the Cloudberry adaptation needs an additional guard to disable pgvector's nested parallel worker path on segments, or to fall back to serial build when running under Cloudberry.
### How to reproduce
## Steps to Reproduce
```sql
CREATE DATABASE cb_pgvector_test;
\c cb_pgvector_test
CREATE EXTENSION vector;
CREATE TABLE vec_parallel_test (
id int,
embedding vector(3)
) DISTRIBUTED BY (id);
INSERT INTO vec_parallel_test
SELECT
i,
format('[%s,%s,%s]', random(), random(), random())::vector
FROM generate_series(1, 100000) AS s(i);
SET client_min_messages = DEBUG;
SET min_parallel_table_scan_size = 1;
SET max_parallel_workers = 8;
SET max_parallel_maintenance_workers = 2;
SET maintenance_work_mem = '512MB';
CREATE INDEX vec_parallel_hnsw_idx
ON vec_parallel_test
USING hnsw (embedding vector_l2_ops);
Actual Behavior
The index build starts, pgvector reports that it is using parallel workers, and then the backend crashes.
Full error log:
DEBUG: Message type M received by from libpq, len = 149 (seg0 127.0.1.1:6000 pid=1047933)
DEBUG: Message type M received by from libpq, len = 149 (seg1 127.0.1.1:6001 pid=1047934)
DEBUG: Calling GetNewRelFileNode returns new relfilenode = 82100
DEBUG: building index "vec_parallel_hnsw_idx" on table "vec_parallel_test" serially
DEBUG: using 2 parallel workers
DEBUG: leader processed 0 tuples
LOG: An exception was encountered during the execution of statement: CREATE INDEX vec_parallel_hnsw_idx
ON vec_parallel_test
USING hnsw (embedding vector_l2_ops);
WARNING: terminating connection because of crash of another server process
DETAIL: The postmaster has commanded this server process to roll back the current transaction and exit, because another server process exited abnormally and possibly corrupted shared memory.
HINT: In a moment you should be able to reconnect to the database and repeat your command.
ERROR: Unexpected internal error (assert.c:48) (assert.c:48)
DETAIL: FailedAssertion("ht->size > 0", File: "udp/ic_udpifc.c", Line: 2013)
CONTEXT: parallel worker
server closed the connection unexpectedly
This probably means the server terminated abnormally
before or while processing the request.
The connection to the server was lost. Attempting reset: Failed.
Operating System
OS: Ubuntu 24.04.3 LTS (Noble Numbat)
Anything else
No response
Are you willing to submit PR?
Code of Conduct
pgvector and Cloudberry Database version
Environment
1 coordinator
2 primary segments
observed segment logs from seg0 127.0.1.1:6000 and seg1 127.0.1.1:6001
psql --version:
psql (Apache Cloudberry) 14.4
pg_config --version:
PostgreSQL 14.4
SELECT version();:
PostgreSQL 14.4 (Apache Cloudberry 3.0.0-devel+dev.2227.g017b06b4f86 build dev) on x86_64-pc-linux-gnu, compiled by gcc-11 (Ubuntu 11.5.0-1ubuntu1~24.04) 11.5.0, 64-bit compiled on Feb 2 2026 05:21:20 (with assert checking)
SHOW server_version;:
14.4
pgvector extension version: 0.8.0
~/projects/pgvector$ git remote -v
origin git@github.com:cloudberry-contrib/pgvector.git (fetch)
origin git@github.com:cloudberry-contrib/pgvector.git (push)
What happened
Summary
When creating a pgvector HNSW index on Apache Cloudberry with parallel maintenance workers enabled, the index build enters the parallel worker path and then crashes the server process.
The log shows pgvector starts using parallel workers:
Actual Behavior
The index build starts, pgvector reports that it is using parallel workers, and then the backend crashes.
Full error log:
Operating System
OS: Ubuntu 24.04.3 LTS (Noble Numbat)
Anything else
No response
Are you willing to submit PR?
Code of Conduct