Blog
What we've learned so far about the KV Store in Splunk 10.6
Splunk 10.6 replaces the MongoDB behind the KV Store with PostgreSQL. The REST API stays the same, and most apps keep working. We tested one of our apps, which keeps all of its state in the KV Store, on Splunk 9.4.15, 10.4.3 and 10.6.0.5, on a single instance and on a search head cluster. These are the five things we have learned so far that admins and app developers should know about.
1. KV Store values end up in _internal by default
This is the one to act on. On 10.6 the KV Store’s new kvstore-pdl service logs its database
operations at INFO, and those lines include the values of the documents it writes. The log is
$SPLUNK_HOME/var/log/splunk/sup-pkg-kvstore-pdl-stdout.log, and Splunk indexes it into
_internal. So anything an app keeps in the KV Store can be read by anyone who can search
_internal, whatever the collection’s own permissions say. _internal is admin-only by default,
but many deployments open it to monitoring or operations roles.
You can check it with a marker value in a throwaway collection:
1curl -k -u admin:changeme \
2 https://localhost:8089/servicesNS/nobody/search/storage/collections/config -d name=notes
3curl -k -u admin:changeme \
4 https://localhost:8089/servicesNS/nobody/search/storage/collections/data/notes \
5 -H 'Content-Type: application/json' -d '{"_key":"n1","secret":"canary-12345"}'
6grep canary-12345 $SPLUNK_HOME/var/log/splunk/sup-pkg-kvstore-pdl-stdout.logSplunk’s supported fix is a logging override. On every search head, and on each member of a
search head cluster, create $SPLUNK_HOME/etc/log-node-platform-local.cfg:
1[kvstore-pdl]
2rootCategory = ERROR,rootAppenderThen restart Splunk. We verified it on Splunk Enterprise 10.6.0.5: writes keep working, and new
values reach neither the log nor _internal. We haven’t tested Splunk Cloud Platform.
If you build apps, don’t count on every admin doing this: treat what you write to the KV Store as
readable by anyone who can search _internal. Keep secrets in storage/passwords, and encrypt
other sensitive values before they reach the KV Store.
2. Rows with a NUL character are dropped by the upgrade
PostgreSQL cannot store the NUL character (U+0000). On 10.6 a KV Store write that contains one
is refused with a 422. The upgrade is worse: the migration silently drops every stored row that
contains a NUL, and still reports success. The only trace is an INFO line with skippedRows.
We haven’t seen a NUL arrive in real data: we found this by writing one on purpose. In our app one could come from text a user pastes into the chat, from search results saved into a conversation, or from a model’s reply. The same goes for any app that stores text it didn’t write.
- Admins: count the rows of the collections that matter before the upgrade, and again after.
- Developers: replace or strip U+0000 before every write, and in queries too.
3. Case-insensitive $regex only folds ASCII
{"$regex": "^überprüfung$", "$options": "i"} finds Überprüfung on 10.4, but not on 10.6:
$options: "i" now only folds A–Z. If your app has a search box, users with names outside
plain ASCII will quietly stop finding things.
The fix that works on every version is to build the folding into the pattern: turn each letter
into a bracket of every character that is the same letter ignoring case ([Üü][Bb][Ee][Rr]…,
and [Σσς] for sigma) and escape everything else. A bracket matches one character, so this
covers one-to-one case pairs only: ß still won’t match SS.
4. Every 404 is now an ERROR in splunkd.log
On 10.6 every 404 from the KV Store is logged at ERROR by a new component, KVServiceClient.
Two common patterns turn this into steady noise in _internal:
- Upsert by “update, and create on 404”. POST to
.../data/<collection>/<key>only updates, so every new row costs a 404. Usebatch_savewith the_keyin the body instead: one request, no 404. - Checking whether a document exists by reading it. Use a find instead, which answers with
an empty list:
.../data/<collection>?query={"_key":"k1"}&limit=1.
If you alert on KV Store errors, add KVServiceClient to what you match: the duplicate-key errors
that used to come from KVStorageProvider now come from it too.
5. On a search head cluster, the KV Store can take a while to come up
On a cluster, the 10.6 KV Store is PostgreSQL with one leader and copies on the other members. A member builds its copy from a backup of the leader, and the backup first waits for a checkpoint on the leader. PostgreSQL spreads a checkpoint out over time, so the more the leader has written since its last one, the longer the wait, up to about 13 minutes.
Right after a cluster forms, the leader has a lot to write. In our tests the last member’s KV
Store read failed until about 14 minutes after the cluster formed, and then recovered on its
own. Don’t restart a member whose KV Store reads failed right after the cluster forms. Give
it a quarter of an hour.
Later the wait can be much shorter: on our quiet cluster, a member whose KV Store data we deleted was ready again in under three minutes.