# Security analytics detector matches \_ws\_ on Text fields but fails on Keywords

**URL:** <https://forum.opensearch.org/t/security-analytics-detector-matches-ws-on-text-fields-but-fails-on-keywords/27866>\
**Category:** Security Analytics\
**Tags:** troubleshoot\
**Created:** [February 23, 2026, 1:01pm UTC](https://forum.opensearch.org/t/security-analytics-detector-matches-ws-on-text-fields-but-fails-on-keywords/27866 "2026-02-23T13:01:48Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![mutant](https://avatars.discourse-cdn.com/v4/letter/m/f4b2a3/32.png) [@mutant](https://forum.opensearch.org/u/mutant)\
**Post date:** [February 23, 2026, 1:01pm UTC](https://forum.opensearch.org/t/security-analytics-detector-matches-ws-on-text-fields-but-fails-on-keywords/27866/1 "2026-02-23T13:01:48Z")

</div>

It came to my attention that the detector was triggering many false positives, and it happened after I changed the index’s text fields to keyword. Upon investigation, found out that the rules were replacing whitespaces with “\_ws\_” escape sequence. For this I created two indexes both with just one attribute. In one index the datatype is keyword and the other is text. A test rule was also created.

Here’s an example of the detection logic in the rule:

```auto
detection:
  condition: Selection_1
  Selection_1:
    companyName|all:
      - microsoft corp

```

Here’s the security analytics generated detection query:

```auto
"query": "companyName: \"microsoft_ws_corp\""

```

As keyword field is not analyzed, my understanding is that the keyword detector wasn’t triggered because “\_ws\_” isn’t present in the ingested document.

```auto
{
  "log.attributes.companyName": "microsoft corp"
}

```

But my text detector worked, think it’s because text fields have analyzers.

to. test my theory about whitespaces, I ingested the following document to the keyword index and a finding was generated.

```auto
{
  "log.attributes.companyName": "microsoft_ws_corp"
}

```

I also queried the exact query string, but no documents were returned fro both the indexes, even the document present in finding wasn’t returned. Maybe the way detectors query the indices are different from what I thought. Anyways that’s a topic for another day.

Shouldn’t opensearch handle the difference between text and keyword in security analytics? I thought the escape sequences are kept in place so that it’s handled in different field types. I also found this exact issue raised in github back in May 2024: [SIGMA rule translation -\> lucene query replaces spaces " " with "\_ws\_" which lucene doesnt understand. · Issue #1024 · opensearch-project/security-analytics · GitHub](https://github.com/opensearch-project/security-analytics/issues/1024) . Someone tried fixing the issue, but they ultimately gave up.

What can be done for the detector to work properly for keyword fields? I know reverting back to text field is an option, but do I have any other options? I explored the usage of custom analyzers, but my application does alot of querying in the indexes, so I fear all that will be affected. Any solution to this? Why was the space replaced with “\_ws\_” which ultimately made the detector to fail for keyword fields?

---

<div class="post-metadata">

**Author:** ![Anthony](https://sea1.discourse-cdn.com/flex019/user_avatar/forum.opensearch.org/anthony/32/9939_2.png) [@Anthony](https://forum.opensearch.org/u/Anthony)\
**Post date:** [March 9, 2026, 1:40pm UTC](https://forum.opensearch.org/t/security-analytics-detector-matches-ws-on-text-fields-but-fails-on-keywords/27866/2 "2026-03-09T13:40:09Z")

</div>

@mutant Thank you for the question. One of the workarounds that you are use is a custom analyzer, see example below:

Create the index with custom analyzer:

```auto
PUT /test-sa-bug
{
  "settings": {
    "analysis": {
      "analyzer": {
        "rule_analyzer": {
          "tokenizer": "keyword",
          "char_filter": ["rule_ws_filter"]
        }
      },
      "char_filter": {
        "rule_ws_filter": {
          "type": "pattern_replace",
          "pattern": "(_ws_)",
          "replacement": " "
        }
      }
    }
  },
  "mappings": {
    "properties": {
      "companyName": {
        "type": "keyword",
        "fields": {
          "text": {
            "type": "text",
            "analyzer": "rule_analyzer"
          }
        }
      }
    }
  }
}

```

Create a document:

```auto
POST /test-sa-bug/_doc
{ "companyName": "microsoft corp" }

```

Use CompanyName.text in query:

```auto
GET /test-sa-bug/_search
{
  "query": {
    "query_string": { "query": "companyName.text:\"microsoft corp\"" }
  }
}

```

If you want this to be applied to existing index, you would need to reindex to a new index with custom analyzer.

Hope this helps

---

<div class="post-metadata">

**Author:** ![mutant](https://avatars.discourse-cdn.com/v4/letter/m/f4b2a3/32.png) [@mutant](https://forum.opensearch.org/u/mutant)\
**Post date:** [March 11, 2026, 6:43am UTC](https://forum.opensearch.org/t/security-analytics-detector-matches-ws-on-text-fields-but-fails-on-keywords/27866/3 "2026-03-11T06:43:10Z")

</div>

Yes, I did try custom analyzers. It worked with the detectors. My application also uses term and wildcard queries to retrieve data from opensearch too, so that also worked. But it didn’t work for aggregation queries, got an illegal argument exception for it. Is there any way to make aggregation queries work with text field, but keyword tokenizer?

Multi-fields seem to be a solid option too, but that’ll increase out storage costs

Sidenote, I also tried using custom normalizers on keyword fields. Unfortunately the detector creation failed via UI. Then I created a detector and then added the index via dev tools, that made the cluster to collapse. Saw lots of alerting exceptions in docker logs. Wasn’t able to recover the cluster, had to clear the volume and start from scratch. So I’m ruling out normalizers.

---

<div class="post-metadata">

**Author:** ![Anthony](https://sea1.discourse-cdn.com/flex019/user_avatar/forum.opensearch.org/anthony/32/9939_2.png) [@Anthony](https://forum.opensearch.org/u/Anthony)\
**Post date:** [March 11, 2026, 12:35pm UTC](https://forum.opensearch.org/t/security-analytics-detector-matches-ws-on-text-fields-but-fails-on-keywords/27866/4 "2026-03-11T12:35:00Z")

</div>

@mutant have you explored using [`fielddata: true`](https://docs.opensearch.org/latest/mappings/supported-field-types/text/#parameters),

You should be able to use the following:

```auto
PUT /test-sa
{
  "settings": {
    "analysis": {
      "analyzer": {
        "rule_analyzer": {
          "tokenizer": "keyword",
          "char_filter": ["rule_ws_filter"]
        }
      },
      "char_filter": {
        "rule_ws_filter": {
          "type": "pattern_replace",
          "pattern": "(_ws_)",
          "replacement": " "
        }
      }
    }
  },
  "mappings": {
    "properties": {
      "companyName": {
        "type": "keyword",
        "fields": {
          "text": {
            "type": "text",
            "analyzer": "rule_analyzer",
            "fielddata": true
          }
        }
      }
    }
  }
}

POST /test-sa/_bulk
{ "index": {} }
{ "companyName": "microsoft corp" }
{ "index": {} }
{ "companyName": "microsoft corp" }
{ "index": {} }
{ "companyName": "apple inc" }

GET /test-sa/_search
{
  "size": 0,
  "aggs": {
    "by_company": {
      "terms": { "field": "companyName.text" }
    }
  }
}

```

This will however increase your heap usage as I think fielddata loads into JVM heap on first aggregation and stays cached there, rather than using disk.

---

<div class="post-metadata">

**Author:** ![mutant](https://avatars.discourse-cdn.com/v4/letter/m/f4b2a3/32.png) [@mutant](https://forum.opensearch.org/u/mutant)\
**Post date:** [March 12, 2026, 4:45am UTC](https://forum.opensearch.org/t/security-analytics-detector-matches-ws-on-text-fields-but-fails-on-keywords/27866/5 "2026-03-12T04:45:14Z")

</div>

fieldData isn’t a risk that I want to take haha.

Anyways can we expect a fix for detectors to work with keyword fields? And, any expected fix for keyword normalizers to work with detectors?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/flex019/uploads/mauve_hedgehog/original/2X/3/33547ea01a5b12dcca2958411d3edd97ae2ea8c1.png) [@system](https://forum.opensearch.org/u/system)\
**Post date:** [May 11, 2026, 4:46am UTC](https://forum.opensearch.org/t/security-analytics-detector-matches-ws-on-text-fields-but-fails-on-keywords/27866/6 "2026-05-11T04:46:06Z")

</div>

This topic was automatically closed 60 days after the last reply. New replies are no longer allowed.
