Shards appear to have negative size, and block listing them

Versions (relevant - OpenSearch/Dashboard/Server OS/Browser): 3.8.0

Describe the issue: Some shards appear to report a negative size. It’s impossible to list the shards because of this.

Configuration: Read/write separated (remote storage is minio), currently 20 data nodes and 5 search nodes. 5 data nodes also are cluster_managers. The index has ZERO REPLICAS.

Relevant Logs or Screenshots:

The problem is that I cannot list shards for the affected index:

$ curl -ksSu "$creds" https://localhost:9200/_list/shards/big_index
{"error":{"root_cause":[{"type":"illegal_argument_exception","reason":"Values less than -1 bytes are not supported: -505064b"}],"type":"illegal_argument_exception","reason":"Values less than -1 bytes are not supported: -505064b"},"status":400}

There is more than one shard with this issue because I also saw these values in the error message while I was adding data nodes:
-943336b
-1001992b
-493792b

From what I can see, everything is working except listing the shards.

How do I fix this? Is there a way to still get the shards listed despite this bug?
I know about cloning my index and deleting the old one (what’s happening now but it will a full day), is there any other way to fix this?

I don’t know when the problem happened, but since the cluster and the index were created, all the operations that were done were:

  • restart one data node
  • cut internet for 6 minutes on one data node
  • restart search nodes
  • transform 10 search nodes (from 15) to data nodes, one by one

I’ll have to redo all this if you need me to try to reproduce.

@Camusensei Can you list shards related to the affected indices? If yes, are they all located on the same node?
Do you see any errors regarding affected indices or it’s shards or disk space watermarks?
Do you monitor disk space on all data nodes?

The error only comes if I try to list all shards or just the affected index’s shards (that’s how I found which index was affected, by trial and error on all indexes until I got the error)
I do monitor disk space on all data nodes, I had no space issue <25% for all data nodes

When I wrote the ticket, my situation cluster looked like this:

20 data (contain the indexes)
5 search (contain search replicas)

After writing this ticket, I excluded the first 10 data nodes, effectively moving all the shards from 20 nodes to the second set of 10 nodes:

10 data (empty)
10 data (contain the indexes)
5 search (contain search replicas)

I reinstalled those 10 empty nodes (that was the operation I needed to do). Then I moved the data back to the first 10 nodes by excluding the second set of 10 nodes:

10 data (contain the indexes)
10 data (empty)
5 search (contain search replicas)

Finally I changed the roles of the empty data nodes back to search nodes:

10 data (contain the indexes)
10 search (empty)
5 search (contain search replicas)

Now I can see that the issue is gone (I am able to list the shards again).
(I also see that I forgot to remove the exclusion which explains why 10 of my search nodes were empty)
Unfortunately I do not know at which step the problem went away as I completely forgot about the issue while I was performing my operation.

But if it happens again, it would be nice to have the code still list the shards that have a positive size despite showing the error… Which would help troubleshoot which shards have a negative size and why.