# Recomendations for k8s logs shipping operator

**URL:** https://forum.opensearch.org/t/recomendations-for-k8s-logs-shipping-operator/11505
**Category:** DevOps
**Created:** [November 15, 2022, 2:09pm UTC](https://forum.opensearch.org/t/recomendations-for-k8s-logs-shipping-operator/11505 "2022-11-15T14:09:41Z")
**Posts on this page:** 5
**Page:** 1

<div class="post-metadata">

### Author: ![tru64jurus](https://avatars.discourse-cdn.com/v4/letter/t/85e7bf/32.png) [@tru64jurus](https://forum.opensearch.org/u/tru64jurus)
#### Post date: [November 15, 2022, 2:09pm UTC](https://forum.opensearch.org/t/recomendations-for-k8s-logs-shipping-operator/11505/1 "2022-11-15T14:09:41Z")

</div>

**Versions** (relevant - OpenSearch/Dashboard/Server OS/Browser):

2.3.0

**Describe the issue** :

Looking for recommendations for Kubernetes logging operator to collect logs from all pods and k8s hosts misc logs. The one we use now runs into performance issues where logs collection stops .

**Configuration** :

Current setup fluentbit → fluentd → kafka \<----\> logstash → Opensearch.

**Relevant Logs or Screenshots** :

---

<div class="post-metadata">

### Author: ![Nicolaegis](https://avatars.discourse-cdn.com/v4/letter/n/d6d6ee/32.png) [@Nicolaegis](https://forum.opensearch.org/u/Nicolaegis)
#### Post date: [November 15, 2022, 3:41pm UTC](https://forum.opensearch.org/t/recomendations-for-k8s-logs-shipping-operator/11505/2 "2022-11-15T15:41:22Z")

</div>

Hi @tru64jurus,

The setup you have looks pretty good and it may need some tuning on k8s and fluentbit side.  
I would like to know more about the performance issues you have.

## Thank you, Nicoale Vartolomei

Elasticsearch/OpenSearch & Solr Consulting, Production Support & Training Sematext Cloud - Full Stack Observability

> **[Sematext | IT System Monitoring Tools for DevOps](https://sematext.com/)**
>
> IT system monitoring and management tools for DevOps who need 24x7 live visibility into their infrastructure.

---

<div class="post-metadata">

### Author: ![tru64jurus](https://avatars.discourse-cdn.com/v4/letter/t/85e7bf/32.png) [@tru64jurus](https://forum.opensearch.org/u/tru64jurus)
#### Post date: [November 15, 2022, 6:28pm UTC](https://forum.opensearch.org/t/recomendations-for-k8s-logs-shipping-operator/11505/3 "2022-11-15T18:28:35Z")

</div>

@Nicolaegis Here is example of issues reported fluentbit not recovering after forward upstream fluentd restart .

> <https://github.com/fluent/fluent-bit/issues/3308>
>
> \## Bug Report
> 
> \*\*Describe the bug\*\*
> Using Fluentbit forwarding into Fluentd u…pstream is working fine but when I restart upstream Fluentd, I will start getting following errors that are fine:
> \`\`\`
> \[2021/03/31 07:48:25\] \[warn\] \[engine\] chunk '7-1617176895.69632323.flb' cannot be retried: task\_id=7, input=forward.0 \> output=forward.0
> \[2021/03/31 07:48:25\] \[error\] \[net\] TCP connection failed: fluentd:32233 (Connection refused)
> \[2021/03/31 07:48:25\] \[error\] \[net\] cannot connect to fluentd:32233
> \[2021/03/31 07:48:25\] \[error\] \[output:forward:forward.0\] no upstream connections available
> \[2021/03/31 07:48:25\] \[warn\] \[engine\] chunk '7-1617176882.905759681.flb' cannot be retried: task\_id=18, input=systemd.1 \> output=forward.0
> \[2021/03/31 07:48:25\] \[error\] \[net\] TCP connection failed: fluentd:32233 (Connection refused)
> \[2021/03/31 07:48:25\] \[error\] \[net\] cannot connect to fluentd:32233
> \[2021/03/31 07:48:25\] \[error\] \[output:forward:forward.0\] no upstream connections available
> \[2021/03/31 07:48:25\] \[warn\] \[engine\] chunk '7-1617176889.866633403.flb' cannot be retried: task\_id=47, input=emitter\_for\_rewrite\_tag.4 \> output=forward.0
> \[2021/03/31 07:48:25\] \[error\] \[net\] TCP connection failed: fluentd:32233 (Connection refused)
> \[2021/03/31 07:48:25\] \[error\] \[net\] cannot connect to fluentd:32233
> \[2021/03/31 07:48:25\] \[error\] \[output:forward:forward.0\] no upstream connections available
> \[2021/03/31 07:48:25\] \[warn\] \[engine\] chunk '7-1617176892.116668507.flb' cannot be retried: task\_id=15, input=forward.0 \> output=forward.0
> \`\`\`
> and then when upstream goes up it will start spamming console forever with connection timed out errors:
> \`\`\`
> \[2021/03/30 21:54:41\] \[error\] \[upstream\] connection #172 to fluentd:32233 timed out after 10 seconds
> \[2021/03/30 21:54:41\] \[error\] \[upstream\] connection #174 to fluentd:32233 timed out after 10 seconds
> \[2021/03/30 21:54:41\] \[error\] \[upstream\] connection #191 to fluentd:32233 timed out after 10 seconds
> \[2021/03/30 21:54:41\] \[error\] \[upstream\] connection #174 to fluentd:32233 timed out after 10 seconds
> \[2021/03/30 21:54:41\] \[error\] \[upstream\] connection #176 to fluentd:32233 timed out after 10 seconds
> \[2021/03/30 21:54:41\] \[error\] \[upstream\] connection #174 to fluentd:32233 timed out after 10 seconds
> \[2021/03/30 21:54:41\] \[error\] \[upstream\] connection #171 to fluentd:32233 timed out after 10 seconds
> \[2021/03/30 21:54:41\] \[error\] \[upstream\] connection #175 to fluentd:32233 timed out after 10 seconds
> \[2021/03/30 21:54:41\] \[error\] \[upstream\] connection #176 to fluentd:32233 timed out after 10 seconds
> \[2021/03/30 21:54:41\] \[error\] \[upstream\] connection #173 to fluentd:32233 timed out after 10 seconds
> \[2021/03/30 21:54:41\] \[error\] \[upstream\] connection #178 to fluentd:32233 timed out after 10 seconds
> \[2021/03/30 21:54:41\] \[error\] \[upstream\] connection #172 to fluentd:32233 timed out after 10 seconds
> 
> \`\`\`
> 
> It basically spams log with 12 messages every second from every affected Fluentbit (there can be hundreds) and it will not recover even after upstream Fluentd is up again. Once fluentbit is restarted, error goes away. I think this is caused by retries of chunks that it failed to submit because I cannot see any successful retries in log.
> 
> Moreover it seems that Prometheus metrics are no longer updated with retry counts. Metric \`fluentbit\_output\_errors\_total{name="forward.0"}\` is still 0, same for \`fluentbit\_output\_retries\_failed\_total{name="forward.0"}\` and \`fluentbit\_output\_retries\_total{name="forward.0"}\`
> 
> I tried to downgrade to Fluentbit 1.6.10 and it works fine so I am suspecting it's caused by multi-worker feature in 1.7.x.
> 
> \*\*To Reproduce\*\*
> 
> \- Steps to reproduce the problem:
> 
> \`\`\`
> output-forward.conf: |
> \[OUTPUT\]
> Name forward
> workers 1
> Match \*
> Self\_Hostname fluentbit
> Host ${FLUENT\_FORWARD\_HOST}
> Port ${FLUENT\_FORWARD\_PORT}
> tls On
> tls.verify On
> tls.ca\_file /secrets/identity/server\_ca.crt
> tls.crt\_file /secrets/identity/client.crt
> tls.key\_file /secrets/identity/client.key
> \`\`\`
> 
> \*\*Expected behavior\*\*
> \- Fluentbit should not spam logs so much when upstream is down.
> \- Fluentbit should recover properly
> \- Metrics should be properly updated to be able to determine wrong behavior in monitoring
> 
> \*\*Your Environment\*\*
> \* Version used: 1.7.2
> \* Configuration:
> \* Environment name and version (e.g. Kubernetes? What version?): Kubernetes
> \* Server type and version:
> \* Operating System and version:
> \* Filters and plugins:
> 
> \*\*Additional context\*\*

> <https://github.com/fluent/fluent-bit/issues/2274>
>
> \## Bug Report
> First of thank you for creating this project and for diligently w…orking to make it better. I am trying to create a log pipeline for our metrics tracking using application logs. I installed fluent-bit on the app server which tails the logs and forwards it (via upstream) to a fluentd box for aggregation.
> 
> \*\*Describe the bug\*\*
> When fluent bit is restarted (or stopped and started), logs are getting lost. I believe the logs that are in the memory buffer are getting dropped when the service is stopped. When the service resumes (restarts) the logs continue forwarding several documents later. By varying the "Mem\_Buf\_Limit" setting I was able to increase/decrease the number of logs lost by the restart.
> 
> \*\*To Reproduce\*\*
> 1. Create a list of json documents with a serially incremented variable.
> \`\`\`
> { mssg: 'the value is ', i: 1000 }
> { mssg: 'the value is ', i: 1001 }
> { mssg: 'the value is ', i: 1002 }
> { mssg: 'the value is ', i: 1003 }
> { mssg: 'the value is ', i: 1004 }
> { mssg: 'the value is ', i: 1005 }
> { mssg: 'the value is ', i: 1006 }
> { mssg: 'the value is ', i: 1007 }
> { mssg: 'the value is ', i: 1008 }
> \`\`\`
> Note: my script created 10M lines to test various features.
> 2. Update the config to tail the file.
> 3. Start the fluent bit service.
> 4. After some time stop the service. (note the last document received by fluentD)
> \`\`\`
> 2020-06-18T16:09:30+00:00 abhi\_event\_manager {"log":"{ mssg: 'the value is ', i: 391336 }","td\_host":"poc-abhi"}
> \`\`\`
> 5. After a few seconds start the service again.
> 6. check the document after the stop, notice the log output has a gap in the series. 
> \`\`\`
> 2020-06-18T16:10:43+00:00 abhi\_event\_manager {"log":"{ mssg: 'the value is ', i: 399626 }","td\_host":"poc-abhi"}
> \`\`\`
> In this case 8290 lines were lost, this number increases as the Mem\_buf\_limit increases.
> \`\`\`
> .
> .
> .
> 2020-06-18T16:09:30+00:00 abhi\_event\_manager {"log":"{ mssg: 'the value is ', i: 391336 }","td\_host":"poc-abhi"}
> 2020-06-18T16:10:43+00:00 abhi\_event\_manager {"log":"{ mssg: 'the value is ', i: 399626 }","td\_host":"poc-abhi"}
> .
> .
> \`\`\`
> 
> 
> \*\*Expected behavior\*\*
> It would be great if the logs are not dropped during restart. (may be the offset stored in the db should track the position of the log that was sent out and not just when it was read into the engine) During restart logs in the buffer should get flushed to the filesystem before shutdown and on startup first process these files before reading new logs.
> 
> \*\*Your Environment\*\*
> \* Version used: 1.3.5 and 1.4.6
> \* Configuration: 
> (Note: I tried change storage.type to memory and filesystem, changing Mem\_Buf\_Limit, Buffer\_Chunk\_Size,Buffer\_Max\_Size but nothing helped)
> \`\`\`
> \[SERVICE\]
> 
> Flush 1
> 
> Daemon Off
> 
> Log\_Level trace
> Log\_File /var/log/td\_bit.log
> 
> Parsers\_File parsers.conf
> Plugins\_File plugins.conf
> 
> HTTP\_Server On
> HTTP\_Listen 0.0.0.0
> HTTP\_Port 2020
> 
> storage.path /var/log/tdbit\_storage/
> storage.sync full
> storage.checksum off
> storage.backlog.mem\_limit 1M
> 
> \[INPUT\]
> Name tail
> Tag abhi\_event\_manager
> Path /var/log/test\_logs/test15.log
> Db /var/log/td.db
> Mem\_Buf\_Limit 1M
> Parser json
> Buffer\_Chunk\_Size 1k
> Buffer\_Max\_Size 1k
> storage.type memory
> 
> \[FILTER\]
> Name record\_modifier
> Match \*
> Record td\_host poc-abhi
> 
> \[OUTPUT\]
> Name forward
> Match abhi\_event\_manager
> Upstream upstream.conf
> Self\_Hostname poc-abhi
> Retry\_Limit False
> \`\`\`
> 
> \* Environment name and version (e.g. Kubernetes? What version?): Installed via rpm (1.3.5) and via build (1.4.6) 
> \* Server type and version: virtual machine on openstack
> \* Operating System and version: Centos 7
> \* Filters and plugins: Filter mentioned above, no plugins.
> 
> On the fluentD side, for testing I am just outputting the logs to a file: (below is the config)
> \`\`\`
> \<match abhi\_event\_manager\>
> @type file
> path /var/log/abhi\_evm
> \</match\>
> \`\`\`
> 
> The goal is to be able to forward logs using fluent bit from the application servers to a centralized fluentD where we would perform aggregation on the log events and use it for metrics reporting.
> So losing logs will lead to inaccurate metrics.

> <https://github.com/fluent/fluent-bit/issues/6010>
>
> \## Bug Report
> 
> \*\*Describe the bug\*\*
> I am running into issues with k8s fluentb…it not recovering after fluentd restart. I have read all similar github issues opened for this error, non of workarounds resolved the issue on my case.
> 
> Setup I have is fluentbit (1.9.7) --\> fluentd --\> kafka(3.2.1) , fluentd service is ClusterIP .
> 
> Appreciate any pointer to remediate this issue.
> 
> Thanks
> 
> Summary of changes tried 
> 
> 
> 
> \*\*To Reproduce\*\*
> \- Restarted fluetnd statfulset 
> \- Example log message if applicable:
> \`\`\`
> Errors I see when fluentbit 
> 
> \[tls\] error: unexpected EOF
> \[output:forward:forward.0\] no upstream connections available
> 
> Then followed by several logs of 
> 
> \[info\] \[task\] re-schedule retry=0x7ff34206b938 2040 in the next 56 seconds
> 
> \`\`\`
> \- Steps to reproduce the problem:
> Restart fluentd statefulset. 
> 
> \*\*Expected behavior\*\*
> 
> fluentbit will test connectivity to fluentd then resume sending logs. 
> 
> 
> \*\*Your Environment\*\*
> 
> \* Version used: 1.9.7
> \* Configuration:
> \* Environment name and version (e.g. Kubernetes? What version?): 1.22
> \* Server type and version: ubuntu 20
> \* Operating System and version:
> \* Filters and plugins:
> 
> \*\*Additional context\*\*
> Summary of changes tried 
> 
> \`\`\`fluentbit 
> net.keepalive\_max\_recycle 100 and 200
> mem\_buf\_limit 5MB and 10MB 20MB
> buffer\_chunk\_size 1M
> buffer\_max\_size 1M and 5MB. 
> 
> fluentd 
> 
> Tried fluentd statefulset of 1 and 2.

---

<div class="post-metadata">

### Author: ![mikeyGlitz](https://avatars.discourse-cdn.com/v4/letter/m/bc8723/32.png) [@mikeyGlitz](https://forum.opensearch.org/u/mikeyGlitz)
#### Post date: [January 7, 2023, 12:34pm UTC](https://forum.opensearch.org/t/recomendations-for-k8s-logs-shipping-operator/11505/4 "2023-01-07T12:34:21Z")

</div>

This logging operator will enable you to push logs into OpenSearch the same way your current configuration works:

fluentbit → fluentd → \<----\> logstash → Opensearch.  
Logging → Flow → output (OpenSearch)

> **[Store NGINX access logs in Elasticsearch with Logging operator · Banzai Cloud](https://banzaicloud.com/docs/one-eye/logging-operator/quickstarts/es-nginx/)**
>
> Bringing cloud native to the enterprise, simplifying the transition to microservices on Kubernetes

---

<div class="post-metadata">

### Author: ![tru64jurus](https://avatars.discourse-cdn.com/v4/letter/t/85e7bf/32.png) [@tru64jurus](https://forum.opensearch.org/u/tru64jurus)
#### Post date: [February 5, 2023, 2:12am UTC](https://forum.opensearch.org/t/recomendations-for-k8s-logs-shipping-operator/11505/5 "2023-02-05T02:12:47Z")

</div>

@mikeyGlitz this logging operator have an issue with fluentbit failing to reconnect to fluentd after fluentd statefulset restart .
