nix-oss opened a new issue, #1885:
URL: https://github.com/apache/cloudberry/issues/1885
### Apache Cloudberry version
2.1.0-incubating
### What happened
Hi team,
When user-defined tablespaces exist in the cluster, gpexpand does not create
the `newTableSpaceInfo.json` file in the new segment template when running with
a config file (non-interactive mode).
This leads to broken symlinks for tablespaces and causes failures during the
prepare_schema() stage.
The problem is present in two methods within the same file:
1. `read_tablespace_file()` in `gpMgmt/bin/gpexpand` (lines 1219–1285)
2. `generate_tablespace_inputfile()` in `gpMgmt/bin/gpexpand` (lines
1114–1148)
- `read_tablespace_file()` (line 1244):
```python
tblspc_oids = os.listdir(coordinator_tblspc_dir)
tblspc_oid_names = self.get_tablespace_oid_names()
flag = False
for oid in tblspc_oids:
if oid in tblspc_oid_names: <---
flag = True
if not flag:
return None
```
- `generate_tablespace_inputfile()` (line 1124)
```python
tblspc_oid_names = self.get_tablespace_oid_names()
tblspc_info = {}
for oid in tblspc_oids:
if oid not in tblspc_oid_names: <---
continue
location =
os.path.dirname(os.readlink(os.path.join(coordinator_tblspc_dir,
oid)))
tblspc_info[oid] = {"location": location,
"name": tblspc_oid_names[int(oid)]}
```
Root cause: type mismatch between string values returned by `os.listdir()`
and integer keys returned by the SQL query `SELECT oid, spcname FROM
pg_tablespace`.
When attaching to the process and debugging, the types differ, causing this
issue:
- type mismatch
<img width="824" height="461" alt="Image"
src="https://github.com/user-attachments/assets/9083b689-c08a-404c-8806-f51f844a2e34"
/>
- `newTableSpaceInfo `is `None`
<img width="599" height="384" alt="Image"
src="https://github.com/user-attachments/assets/9f1077ee-3902-440f-b50a-2fa59862159a"
/>
Impact:
- `newTableSpaceInfo `is always None.
- `_handle_tablespace_template()` is never invoked.
- `newTableSpaceInfo.json` is not included in the template tar.
- `gpconfigurenewsegment `on new hosts does not fix tablespace symlinks.
- new segments start with broken symlinks.
### What you think should happen instead
To fix the issue, unify the data types in both methods:
- In `generate_tablespace_inputfile()`:
```python
for oid in tblspc_oids:
if int(oid) not in tblspc_oid_names:
```
- In `read_tablespace_file()`:
```python
tblspc_oids = os.listdir(coordinator_tblspc_dir)
tblspc_oid_names = self.get_tablespace_oid_names()
flag = False
for oid in tblspc_oids:
if int(oid) in tblspc_oid_names:
```
After this change, `newTableSpaceInfo `is no longer `None`:
<img width="1007" height="631" alt="Image"
src="https://github.com/user-attachments/assets/76009602-1e16-4daf-b0b9-c8b736c918e1"
/>
### How to reproduce
1. Connect to the database:
```sh
gpadmin@cbdb-mdw:~$ psql warehouse
```
2. Create a user-defined tablespace:
```sql
warehouse=# CREATE TABLESPACE ts_stage_logs LOCATION '/tblspc_stage_logs';
```
```
warehouse=# SELECT * FROM pg_tablespace;
oid | spcname | spcowner | spcacl | spcoptions | spcfilehandlersrc
| spcfilehandlerbin
-------+---------------+----------+--------+------------+-------------------+-------------------
1663 | pg_default | 10 | | |
|
1664 | pg_global | 10 | | |
|
17019 | ts_stage_logs | 10 | | |
|
(3 rows)
```
3. Create an append-optimized columnar table and populate it with test data:
```sql
warehouse=# CREATE TABLE logs_aot (
id BIGSERIAL,
log_timestamp TIMESTAMP WITH TIME ZONE DEFAULT NOW(),
log_level VARCHAR(10)
)
WITH (
APPENDONLY = TRUE,
ORIENTATION = COLUMN
)
TABLESPACE ts_stage_logs
DISTRIBUTED BY (id);
warehouse=# INSERT INTO logs_aot (
log_timestamp,
log_level
)
SELECT
NOW() - (random() * INTERVAL '90 days'),
(ARRAY['INFO', 'WARN', 'ERROR', 'DEBUG'])[floor(random() * 4 + 1)]
FROM generate_series(1, 8000);
```
4. Prepare the expansion configuration file for adding 2 new hosts to the
cluster:
```
sdw3-new|sdw3-new|6000|/primary_data/gpseg4|10|4|p
sdw4-new|sdw4-new|7000|/mirror_data/gpseg4|11|4|m
sdw3-new|sdw3-new|6001|/primary_data/gpseg5|12|5|p
sdw4-new|sdw4-new|7001|/mirror_data/gpseg5|13|5|m
sdw4-new|sdw4-new|6000|/primary_data/gpseg6|14|6|p
sdw3-new|sdw3-new|7000|/mirror_data/gpseg6|15|6|m
sdw4-new|sdw4-new|6001|/primary_data/gpseg7|16|7|p
sdw3-new|sdw3-new|7001|/mirror_data/gpseg7|17|7|m
```
5. Start the expansion process (non-interactive mode):
```sh
gpadmin@cbdb-mdw:~$ gpexpand -i expand.cfg
```
6. The following error appears in the logs (log truncated for readability.)
```
20260805:18:33:04:052024 gpexpand:cbdb-mdw:gpadmin-[INFO]:-Heap checksum
setting consistent across cluster
20260805:18:33:04:052024 gpexpand:cbdb-mdw:gpadmin-[INFO]:-Syncing Apache
Cloudberry extensions
20260805:18:33:04:052024 gpexpand:cbdb-mdw:gpadmin-[INFO]:-Locking catalog
20260805:18:33:05:052024 gpexpand:cbdb-mdw:gpadmin-[INFO]:-Locked catalog
20260805:18:33:06:052024 gpexpand:cbdb-mdw:gpadmin-[INFO]:-Creating segment
template
20260805:18:33:07:052024 gpexpand:cbdb-mdw:gpadmin-[INFO]:-Copying
postgresql.conf from existing segment into template
20260805:18:33:08:052024 gpexpand:cbdb-mdw:gpadmin-[INFO]:-Copying
pg_hba.conf from existing segment into template
20260805:18:33:09:052024 gpexpand:cbdb-mdw:gpadmin-[INFO]:-Creating schema
tar file
...
20260805:18:33:26:052024 gpexpand:cbdb-mdw:gpadmin-[INFO]:-Populating
gpexpand.status_detail with data from database postgres
20260805:18:33:27:052024 gpexpand:cbdb-mdw:gpadmin-[INFO]:-Populating
gpexpand.status_detail with data from database warehouse
20260805:18:33:27:052024 gpexpand:cbdb-mdw:gpadmin-[ERROR]:-gpexpand failed:
ERROR: could not open file "pg_tblspc/17019/GPDB_3_302606111/17018/16386": No
such file or directory (seg6 203.0.113.5:6000 pid=21449)
Exiting...
20260805:18:33:27:052024 gpexpand:cbdb-mdw:gpadmin-[ERROR]:-gpexpand is past
the point of rollback. Any remaining issues must be addressed outside of
gpexpand.
20260805:18:33:27:052024 gpexpand:cbdb-mdw:gpadmin-[INFO]:-Shutting down
gpexpand...
```
Hope this analysis is helpful.
### Operating System
Ubuntu 22.04
### Anything else
_No response_
### Are you willing to submit PR?
- [ ] Yes, I am willing to submit a PR!
### Code of Conduct
- [x] I agree to follow this project's [Code of
Conduct](https://github.com/apache/cloudberry/blob/main/CODE_OF_CONDUCT.md).
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]