nix-oss opened a new issue, #1885:
URL: https://github.com/apache/cloudberry/issues/1885

   ### Apache Cloudberry version
   
   2.1.0-incubating
   
   ### What happened
   
   Hi team,
   
   When user-defined tablespaces exist in the cluster, gpexpand does not create 
the `newTableSpaceInfo.json` file in the new segment template when running with 
a config file (non-interactive mode).
   
   This leads to broken symlinks for tablespaces and causes failures during the 
prepare_schema() stage.
   
   The problem is present in two methods within the same file:
   1. `read_tablespace_file()` in `gpMgmt/bin/gpexpand` (lines 1219–1285)
   2. `generate_tablespace_inputfile()` in `gpMgmt/bin/gpexpand` (lines 
1114–1148)
   
   - `read_tablespace_file()` (line 1244):
   ```python
   tblspc_oids = os.listdir(coordinator_tblspc_dir)
   tblspc_oid_names = self.get_tablespace_oid_names()
   flag = False
   for oid in tblspc_oids:
       if oid in tblspc_oid_names:  <---
           flag = True
   if not flag:
       return None  
   ```
   
   - `generate_tablespace_inputfile()` (line 1124)
   ```python
   tblspc_oid_names = self.get_tablespace_oid_names()
   tblspc_info = {}
   
   for oid in tblspc_oids:
       if oid not in tblspc_oid_names:  <---
           continue
       location = 
os.path.dirname(os.readlink(os.path.join(coordinator_tblspc_dir,
                                                           oid)))
       tblspc_info[oid] = {"location": location,
                           "name": tblspc_oid_names[int(oid)]}
   ```
   
   Root cause: type mismatch between string values returned by `os.listdir()` 
and integer keys returned by the SQL query `SELECT oid, spcname FROM 
pg_tablespace`.
   
   When attaching to the process and debugging, the types differ, causing this 
issue:
   
   - type mismatch
   
   <img width="824" height="461" alt="Image" 
src="https://github.com/user-attachments/assets/9083b689-c08a-404c-8806-f51f844a2e34";
 />
   
   - `newTableSpaceInfo `is `None`
   
   <img width="599" height="384" alt="Image" 
src="https://github.com/user-attachments/assets/9f1077ee-3902-440f-b50a-2fa59862159a";
 />
   
   Impact:
   - `newTableSpaceInfo `is always None.
   - `_handle_tablespace_template()` is never invoked.
   - `newTableSpaceInfo.json` is not included in the template tar.
   - `gpconfigurenewsegment `on new hosts does not fix tablespace symlinks.
   - new segments start with broken symlinks.
   
   
   ### What you think should happen instead
   
   To fix the issue, unify the data types in both methods:
   - In `generate_tablespace_inputfile()`:
   ```python
   for oid in tblspc_oids:
     if int(oid) not in tblspc_oid_names:
   ```
   
   - In `read_tablespace_file()`: 
   ```python
   tblspc_oids = os.listdir(coordinator_tblspc_dir)
   tblspc_oid_names = self.get_tablespace_oid_names()
   flag = False
   for oid in tblspc_oids:
     if int(oid) in tblspc_oid_names:
   ```
   
   After this change, `newTableSpaceInfo `is no longer `None`:
   
   <img width="1007" height="631" alt="Image" 
src="https://github.com/user-attachments/assets/76009602-1e16-4daf-b0b9-c8b736c918e1";
 />
   
   
   ### How to reproduce
   
   1. Connect to the database:
   ```sh
   gpadmin@cbdb-mdw:~$ psql warehouse
   ```
   
   2. Create a user-defined tablespace:
   ```sql
   warehouse=# CREATE TABLESPACE ts_stage_logs LOCATION '/tblspc_stage_logs';
   ```
   
   ```
   warehouse=# SELECT * FROM pg_tablespace;
     oid  |    spcname    | spcowner | spcacl | spcoptions | spcfilehandlersrc 
| spcfilehandlerbin 
   
-------+---------------+----------+--------+------------+-------------------+-------------------
     1663 | pg_default    |       10 |        |            |                   
| 
     1664 | pg_global     |       10 |        |            |                   
| 
    17019 | ts_stage_logs |       10 |        |            |                   
| 
   (3 rows)
   ```
   
   3. Create an append-optimized columnar table and populate it with test data:
   ```sql
    warehouse=# CREATE TABLE logs_aot (
        id              BIGSERIAL,
        log_timestamp   TIMESTAMP WITH TIME ZONE DEFAULT NOW(),
        log_level       VARCHAR(10)
    )
    WITH (
        APPENDONLY = TRUE,
        ORIENTATION = COLUMN
    )
    TABLESPACE ts_stage_logs
    DISTRIBUTED BY (id);
   
    warehouse=# INSERT INTO logs_aot (
        log_timestamp,
        log_level
    )
    SELECT
        NOW() - (random() * INTERVAL '90 days'),
        (ARRAY['INFO', 'WARN', 'ERROR', 'DEBUG'])[floor(random() * 4 + 1)]
    FROM generate_series(1, 8000);
   ```
   
   4. Prepare the expansion configuration file for adding 2 new hosts to the 
cluster:
   ```
   sdw3-new|sdw3-new|6000|/primary_data/gpseg4|10|4|p
   sdw4-new|sdw4-new|7000|/mirror_data/gpseg4|11|4|m
   sdw3-new|sdw3-new|6001|/primary_data/gpseg5|12|5|p
   sdw4-new|sdw4-new|7001|/mirror_data/gpseg5|13|5|m
   sdw4-new|sdw4-new|6000|/primary_data/gpseg6|14|6|p
   sdw3-new|sdw3-new|7000|/mirror_data/gpseg6|15|6|m
   sdw4-new|sdw4-new|6001|/primary_data/gpseg7|16|7|p
   sdw3-new|sdw3-new|7001|/mirror_data/gpseg7|17|7|m
   ```
   
   5. Start the expansion process (non-interactive mode):
   ```sh
   gpadmin@cbdb-mdw:~$ gpexpand -i expand.cfg
   ```
   
   6. The following error appears in the logs (log truncated for readability.)
   ```
   20260805:18:33:04:052024 gpexpand:cbdb-mdw:gpadmin-[INFO]:-Heap checksum 
setting consistent across cluster
   20260805:18:33:04:052024 gpexpand:cbdb-mdw:gpadmin-[INFO]:-Syncing Apache 
Cloudberry extensions
   20260805:18:33:04:052024 gpexpand:cbdb-mdw:gpadmin-[INFO]:-Locking catalog
   20260805:18:33:05:052024 gpexpand:cbdb-mdw:gpadmin-[INFO]:-Locked catalog
   20260805:18:33:06:052024 gpexpand:cbdb-mdw:gpadmin-[INFO]:-Creating segment 
template
   20260805:18:33:07:052024 gpexpand:cbdb-mdw:gpadmin-[INFO]:-Copying 
postgresql.conf from existing segment into template
   20260805:18:33:08:052024 gpexpand:cbdb-mdw:gpadmin-[INFO]:-Copying 
pg_hba.conf from existing segment into template
   20260805:18:33:09:052024 gpexpand:cbdb-mdw:gpadmin-[INFO]:-Creating schema 
tar file
   ...
   20260805:18:33:26:052024 gpexpand:cbdb-mdw:gpadmin-[INFO]:-Populating 
gpexpand.status_detail with data from database postgres
   20260805:18:33:27:052024 gpexpand:cbdb-mdw:gpadmin-[INFO]:-Populating 
gpexpand.status_detail with data from database warehouse
   20260805:18:33:27:052024 gpexpand:cbdb-mdw:gpadmin-[ERROR]:-gpexpand failed: 
ERROR:  could not open file "pg_tblspc/17019/GPDB_3_302606111/17018/16386": No 
such file or directory  (seg6 203.0.113.5:6000 pid=21449)
   
   Exiting...
   20260805:18:33:27:052024 gpexpand:cbdb-mdw:gpadmin-[ERROR]:-gpexpand is past 
the point of rollback. Any remaining issues must be addressed outside of 
gpexpand.
   20260805:18:33:27:052024 gpexpand:cbdb-mdw:gpadmin-[INFO]:-Shutting down 
gpexpand...
   ```
   
   Hope this analysis is helpful.
   
   ### Operating System
   
   Ubuntu 22.04
   
   ### Anything else
   
   _No response_
   
   ### Are you willing to submit PR?
   
   - [ ] Yes, I am willing to submit a PR!
   
   ### Code of Conduct
   
   - [x] I agree to follow this project's [Code of 
Conduct](https://github.com/apache/cloudberry/blob/main/CODE_OF_CONDUCT.md).
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to