EKS Cluster's node down

0

We are using EKS and one of our dedicated worker node for TimescaleDB (TSDB) went down. The node soon displayed unreachable taints and a new node got scheduled in place of it soon after.

We want to investigate further as to why the node went down. We have an idea that the TSDB pod was operating on high memory moments before the crash but we would like to be concretely point out the reason for the issue in order to fault-proof it for future.

Can someone suggest a direction to take? We already looked at the logs for the instance that went down, there was nothing that concretely points to a crash or node being unavailable.

Arisht
已提问 6 个月前155 查看次数
没有答案

您未登录。 登录 发布回答。

一个好的回答可以清楚地解答问题和提供建设性反馈,并能促进提问者的职业发展。

回答问题的准则