Blog

What we've learned so far about the KV Store in Splunk 10.6

4 min read Back to all posts
splunk splunk-10.6 kv-store postgresql upgrade search-head-cluster splunk-apps

Splunk 10.6 replaces the MongoDB behind the KV Store with PostgreSQL. The REST API stays the same, and most apps keep working. We tested one of our apps, which keeps all of its state in the KV Store, on Splunk 9.4.15, 10.4.3 and 10.6.0.5, on a single instance and on a search head cluster. These are the five things we have learned so far that admins and app developers should know about.

1. KV Store values end up in _internal by default

This is the one to act on. On 10.6 the KV Store’s new kvstore-pdl service logs its database operations at INFO, and those lines include the values of the documents it writes. The log is $SPLUNK_HOME/var/log/splunk/sup-pkg-kvstore-pdl-stdout.log, and Splunk indexes it into _internal. So anything an app keeps in the KV Store can be read by anyone who can search _internal, whatever the collection’s own permissions say. _internal is admin-only by default, but many deployments open it to monitoring or operations roles.

You can check it with a marker value in a throwaway collection:

bash
1curl -k -u admin:changeme \
2  https://localhost:8089/servicesNS/nobody/search/storage/collections/config -d name=notes
3curl -k -u admin:changeme \
4  https://localhost:8089/servicesNS/nobody/search/storage/collections/data/notes \
5  -H 'Content-Type: application/json' -d '{"_key":"n1","secret":"canary-12345"}'
6grep canary-12345 $SPLUNK_HOME/var/log/splunk/sup-pkg-kvstore-pdl-stdout.log

Splunk’s supported fix is a logging override. On every search head, and on each member of a search head cluster, create $SPLUNK_HOME/etc/log-node-platform-local.cfg:

ini
1[kvstore-pdl]
2rootCategory = ERROR,rootAppender

Then restart Splunk. We verified it on Splunk Enterprise 10.6.0.5: writes keep working, and new values reach neither the log nor _internal. We haven’t tested Splunk Cloud Platform.

If you build apps, don’t count on every admin doing this: treat what you write to the KV Store as readable by anyone who can search _internal. Keep secrets in storage/passwords, and encrypt other sensitive values before they reach the KV Store.

2. Rows with a NUL character are dropped by the upgrade

PostgreSQL cannot store the NUL character (U+0000). On 10.6 a KV Store write that contains one is refused with a 422. The upgrade is worse: the migration silently drops every stored row that contains a NUL, and still reports success. The only trace is an INFO line with skippedRows.

We haven’t seen a NUL arrive in real data: we found this by writing one on purpose. In our app one could come from text a user pastes into the chat, from search results saved into a conversation, or from a model’s reply. The same goes for any app that stores text it didn’t write.

  • Admins: count the rows of the collections that matter before the upgrade, and again after.
  • Developers: replace or strip U+0000 before every write, and in queries too.

3. Case-insensitive $regex only folds ASCII

{"$regex": "^überprüfung$", "$options": "i"} finds Überprüfung on 10.4, but not on 10.6: $options: "i" now only folds A–Z. If your app has a search box, users with names outside plain ASCII will quietly stop finding things.

The fix that works on every version is to build the folding into the pattern: turn each letter into a bracket of every character that is the same letter ignoring case ([Üü][Bb][Ee][Rr]…, and [Σσς] for sigma) and escape everything else. A bracket matches one character, so this covers one-to-one case pairs only: ß still won’t match SS.

4. Every 404 is now an ERROR in splunkd.log

On 10.6 every 404 from the KV Store is logged at ERROR by a new component, KVServiceClient. Two common patterns turn this into steady noise in _internal:

  • Upsert by “update, and create on 404”. POST to .../data/<collection>/<key> only updates, so every new row costs a 404. Use batch_save with the _key in the body instead: one request, no 404.
  • Checking whether a document exists by reading it. Use a find instead, which answers with an empty list: .../data/<collection>?query={"_key":"k1"}&limit=1.

If you alert on KV Store errors, add KVServiceClient to what you match: the duplicate-key errors that used to come from KVStorageProvider now come from it too.

5. On a search head cluster, the KV Store can take a while to come up

On a cluster, the 10.6 KV Store is PostgreSQL with one leader and copies on the other members. A member builds its copy from a backup of the leader, and the backup first waits for a checkpoint on the leader. PostgreSQL spreads a checkpoint out over time, so the more the leader has written since its last one, the longer the wait, up to about 13 minutes.

Right after a cluster forms, the leader has a lot to write. In our tests the last member’s KV Store read failed until about 14 minutes after the cluster formed, and then recovered on its own. Don’t restart a member whose KV Store reads failed right after the cluster forms. Give it a quarter of an hour.

Later the wait can be much shorter: on our quiet cluster, a member whose KV Store data we deleted was ready again in under three minutes.

About Outcold Solutions

Outcold Solutions builds applications for Splunk Enterprise and Splunk Cloud. Our certified monitoring solutions, powered by Collectord, bring logs, metrics and events from Kubernetes, OpenShift and Docker clusters, Linux hosts and Windows containers into Splunk, with the dashboards and alerts that help developers watch their applications and operators keep their clusters healthy. Our search apps query Kubernetes and AWS live from the search bar, with nothing to ingest, and OS AI Agent puts the model you choose to work inside Splunk, with each user's own permissions. Since 2017 we have been helping businesses keep what they need to answer complex questions about their infrastructure in one place.

Red Hat
Splunk
AWS