Skip to content

part: allocate watches fully lazily - #198

Merged
joamaki merged 3 commits into
mainfrom
pr/joamaki/improve-watches
Sep 14, 2026
Merged

joamaki merged 3 commits into
mainfrom
pr/joamaki/improve-watches

Conversation

@joamaki

@joamaki joamaki commented Sep 7, 2026 •

Copy link
Copy Markdown
Contributor

The watch state was still allocated eagerly. This refactores the code to fully allocate watches lazily. The nodes now contain a 'lazyWatchChannel' which is an atomic pointer to 'atomicWatchPointer', so essentially each node has 'atomic.Pointer[atomic.Pointer[*runtime.hchan]]'. The indirection is required as a logical watch may be shared when the tree forks. If we'd just store 'atomic.Pointer[*runtime.hchan]' we wouldn't be able to avoid a double-close when the tree forks as the two instances would not be able to synchronize.

goos: linux
goarch: arm64
pkg: github.com/cilium/statedb/part
              │  before   │              after              │
              │  sec/op   │  sec/op     vs base             │
_Insert          112.1µ      111.1µ            ~ (p=0.280 n=10)
_WatchReplace   108.39µ      88.58µ      -18.27% (p=0.001 n=10)
_GetWatch        17.91µ      19.75µ      +10.28% (p=0.001 n=10)

              │  before   │              after              │
              │   B/op    │   B/op      vs base             │
_Insert         82.28Ki     74.31Ki       -9.68% (p=0.000 n=10)
_WatchReplace   63.81Ki     55.96Ki      -12.30% (p=0.000 n=10)
_GetWatch         0.000       0.000            ~ (p=1.000 n=10)

              │  before   │              after              │
              │ allocs/op │ allocs/op   vs base             │
_Insert          3.067k      2.047k      -33.26% (p=0.000 n=10)
_WatchReplace    2.011k      1.006k      -49.98% (p=0.000 n=10)
_GetWatch         0.000       0.000            ~ (p=1.000 n=10)

Reconciler benchmark (best of 3) before:

1000000 objects reconciled in 1.38 seconds (batch size 1000) 
Throughput 727151.51 objects per second
568MB total allocated, 6015247 in-use objects, 239MB bytes in use

After:

1000000 objects reconciled in 1.34 seconds (batch size 1000) 
Throughput 748576.38 objects per second
552MB total allocated, 5010271 in-use objects, 231MB bytes in use

An isolated comparison of the pointer wrapper with the direct channel pointer shows first channel creation improving from 51.44ns to 42.41ns (-17.54%), 128B to 120B (-6.25%), and 3 to 2 allocations (-33.33%). Accessing an existing channel changes from 1.999ns to 2.042ns (+2.15%).

I also compared against an equivalent atomic.Value implementation. atomic.Value was 31.98% slower when creating the first channel, 6.48% slower when loading an existing channel, 58.14% slower when closing before observation, and 31.54% slower for channel creation followed by close. Allocation counts were equal, but atomic.Value used 128B instead of 120B for channel creation and 16B instead of 8B for close before observation. So I would argue the trade-off of using unsafe.Pointer to grab the *hchan is acceptable. Yes it means StateDB won't work on Go implementations/architectures where chan T is not a pointer, but considering the other uses of unsafe.Pointer in part/node.go this isn't really an issue. It would also mean that if Go ever changes the internal representation things would break here, but tests will easily catch that if it were to happen and we can then change to atomic.Value.

After this PR the part.RootOnlyWatch option doesn't help nearly as much anymore. However removing it completely did increase allocations from 552MB to 581MB in the reconciler benchmark without impact on throughput, so I'll leave it around for now. We can consider removing it in the future to simplify the implementation.

AIL:3

@joamaki
joamaki requested a review from a team as a code owner September 7, 2026 07:31
@joamaki
joamaki requested review from derailed and removed request for a team September 7, 2026 07:31
@github-actions

github-actions Bot commented Sep 7, 2026 •

Copy link
Copy Markdown
$ make
go build ./...
go: downloading go.yaml.in/yaml/v3 v3.0.4
go: downloading github.com/cilium/hive v1.0.4
go: downloading golang.org/x/time v0.15.0
go: downloading github.com/spf13/cobra v1.10.2
go: downloading github.com/spf13/pflag v1.0.10
go: downloading github.com/cilium/stream v0.0.1
go: downloading github.com/liggitt/tabwriter v0.0.0-20181228230101-89fcab3d43de
go: downloading github.com/spf13/viper v1.18.2
go: downloading go.uber.org/dig v1.17.1
go: downloading golang.org/x/term v0.16.0
go: downloading github.com/davecgh/go-spew v1.1.2-0.20180830191138-d8f796af33cc
go: downloading github.com/mitchellh/mapstructure v1.5.0
go: downloading golang.org/x/sys v0.17.0
go: downloading golang.org/x/tools v0.17.0
go: downloading github.com/spf13/cast v1.6.0
go: downloading github.com/fsnotify/fsnotify v1.7.0
go: downloading github.com/sagikazarmark/slog-shim v0.1.0
go: downloading github.com/spf13/afero v1.11.0
go: downloading github.com/subosito/gotenv v1.6.0
go: downloading github.com/hashicorp/hcl v1.0.0
go: downloading gopkg.in/ini.v1 v1.67.0
go: downloading github.com/magiconair/properties v1.8.7
go: downloading github.com/pelletier/go-toml/v2 v2.1.0
go: downloading gopkg.in/yaml.v3 v3.0.1
go: downloading golang.org/x/text v0.14.0
STATEDB_VALIDATE=1 go test ./... -cover -vet=all -test.count 1
go: downloading github.com/stretchr/testify v1.11.1
go: downloading go.uber.org/goleak v1.3.0
go: downloading golang.org/x/exp v0.0.0-20240119083558-1b970713d09a
go: downloading github.com/pmezard/go-difflib v1.0.1-0.20181226105442-5d4384ee4fb2
ok  	github.com/cilium/statedb	330.469s	coverage: 79.6% of statements
ok  	github.com/cilium/statedb/index	0.006s	coverage: 48.6% of statements
ok  	github.com/cilium/statedb/internal	0.015s	coverage: 45.2% of statements
ok  	github.com/cilium/statedb/lpm	5.016s	coverage: 76.1% of statements
ok  	github.com/cilium/statedb/part	84.567s	coverage: 86.8% of statements
ok  	github.com/cilium/statedb/reconciler	0.288s	coverage: 93.3% of statements
	github.com/cilium/statedb/reconciler/benchmark		coverage: 0.0% of statements
	github.com/cilium/statedb/reconciler/example		coverage: 0.0% of statements
go test -race ./... -test.count 1
ok  	github.com/cilium/statedb	46.011s
ok  	github.com/cilium/statedb/index	1.017s
ok  	github.com/cilium/statedb/internal	1.026s
ok  	github.com/cilium/statedb/lpm	2.962s
ok  	github.com/cilium/statedb/part	39.300s
ok  	github.com/cilium/statedb/reconciler	1.355s
?   	github.com/cilium/statedb/reconciler/benchmark	[no test files]
?   	github.com/cilium/statedb/reconciler/example	[no test files]
go test ./... -bench . -benchmem -test.run xxx
goos: linux
goarch: amd64
pkg: github.com/cilium/statedb
cpu: AMD EPYC 7763 64-Core Processor                
BenchmarkDB_WriteTxn_1-4                      	  857932	      1404 ns/op	    712361 objects/sec	     656 B/op	      13 allocs/op
BenchmarkDB_WriteTxn_10-4                     	 1938312	       616.3 ns/op	   1622627 objects/sec	     352 B/op	       6 allocs/op
BenchmarkDB_WriteTxn_100-4                    	 2397102	       524.7 ns/op	   1905800 objects/sec	     317 B/op	       5 allocs/op
BenchmarkDB_WriteTxn_1000-4                   	 2197074	       545.3 ns/op	   1833955 objects/sec	     318 B/op	       5 allocs/op
BenchmarkDB_WriteTxn_100_SecondaryIndex-4     	 1507497	       797.7 ns/op	   1253556 objects/sec	     419 B/op	       7 allocs/op
BenchmarkDB_WriteTxn_CommitOnly_100Tables-4   	 1000000	      1151 ns/op	    1112 B/op	       4 allocs/op
BenchmarkDB_WriteTxn_CommitOnly_1Table-4      	 1689145	       707.9 ns/op	     224 B/op	       4 allocs/op
BenchmarkDB_NewWriteTxn-4                     	 1941988	       616.7 ns/op	     200 B/op	       3 allocs/op
BenchmarkDB_WriteTxnCommit100-4               	  996501	      1177 ns/op	    1096 B/op	       4 allocs/op
BenchmarkDB_NewReadTxn-4                      	641424115	         1.871 ns/op	       0 B/op	       0 allocs/op
BenchmarkDB_Modify-4                          	    1986	    588241 ns/op	   1699984 objects/sec	  342509 B/op	    6073 allocs/op
BenchmarkDB_GetInsert-4                       	    1771	    672203 ns/op	   1487646 objects/sec	  326507 B/op	    6073 allocs/op
BenchmarkDB_RandomInsert-4                    	    2160	    553523 ns/op	   1806609 objects/sec	  318502 B/op	    5073 allocs/op
BenchmarkDB_RandomReplace-4                   	    1130	   1031842 ns/op	    969140 objects/sec	  432325 B/op	    8089 allocs/op
BenchmarkDB_SequentialInsert-4                	    2240	    535663 ns/op	   1866844 objects/sec	  318501 B/op	    5073 allocs/op
BenchmarkDB_SequentialInsert_Prefix-4         	     505	   2342404 ns/op	    426912 objects/sec	 2794465 B/op	   41797 allocs/op
BenchmarkDB_Changes_Baseline-4                	    1861	    628909 ns/op	   1590054 objects/sec	  388270 B/op	    6164 allocs/op
BenchmarkDB_Changes-4                         	     987	   1235749 ns/op	    809226 objects/sec	  604464 B/op	    9332 allocs/op
BenchmarkDB_RandomLookup-4                    	   28449	     42261 ns/op	  23662614 objects/sec	       0 B/op	       0 allocs/op
BenchmarkDB_SequentialLookup-4                	   35428	     33835 ns/op	  29554955 objects/sec	       0 B/op	       0 allocs/op
BenchmarkDB_Prefix_SecondaryIndex-4           	    6830	    167679 ns/op	   5963785 objects/sec	  124952 B/op	    1026 allocs/op
BenchmarkDB_FullIteration_All-4               	     410	   2811226 ns/op	  35571672 objects/sec	     104 B/op	       4 allocs/op
BenchmarkDB_FullIteration_Prefix-4            	     662	   1796876 ns/op	  55652150 objects/sec	     136 B/op	       5 allocs/op
BenchmarkDB_FullIteration_Get-4               	     183	   6273493 ns/op	  15940082 objects/sec	       0 B/op	       0 allocs/op
BenchmarkDB_FullIteration_Get_Secondary-4     	      87	  13490644 ns/op	   7412545 objects/sec	       0 B/op	       0 allocs/op
BenchmarkDB_FullIteration_ReadTxnGet-4        	     178	   6582023 ns/op	  15192898 objects/sec	       0 B/op	       0 allocs/op
BenchmarkDB_PropagationDelay-4                	  663037	      1568 ns/op	        13.00 50th_µs	        16.00 90th_µs	        43.00 99th_µs	     806 B/op	      15 allocs/op
BenchmarkDB_WriteTxn_100_LPMIndex-4           	  524002	      2368 ns/op	    422276 objects/sec	    1574 B/op	      35 allocs/op
BenchmarkDB_WriteTxn_1_LPMIndex-4             	  128254	     14644 ns/op	     68287 objects/sec	   13502 B/op	      76 allocs/op
BenchmarkDB_LPMIndex_Get-4                    	     222	   5525303 ns/op	   1809856 objects/sec	       0 B/op	       0 allocs/op
BenchmarkWatchSet_4-4                         	 2316702	       511.4 ns/op	     296 B/op	       4 allocs/op
BenchmarkWatchSet_16-4                        	  763909	      1559 ns/op	    1096 B/op	       5 allocs/op
BenchmarkWatchSet_128-4                       	   86460	     13720 ns/op	    8904 B/op	       5 allocs/op
BenchmarkWatchSet_1024-4                      	    8307	    136279 ns/op	   73744 B/op	       5 allocs/op
PASS
ok  	github.com/cilium/statedb	44.333s
PASS
ok  	github.com/cilium/statedb/index	0.004s
goos: linux
goarch: amd64
pkg: github.com/cilium/statedb/internal
cpu: AMD EPYC 7763 64-Core Processor                
Benchmark_SortableMutex-4   	 6292402	       190.9 ns/op	       0 B/op	       0 allocs/op
PASS
ok  	github.com/cilium/statedb/internal	1.205s
goos: linux
goarch: amd64
pkg: github.com/cilium/statedb/lpm
cpu: AMD EPYC 7763 64-Core Processor                
Benchmark_txn_insert/batchSize=1-4         	    1830	    644031 ns/op	   1552721 objects/sec	  822405 B/op	   13975 allocs/op
Benchmark_txn_insert/batchSize=10-4        	    3128	    377463 ns/op	   2649267 objects/sec	  369194 B/op	    6668 allocs/op
Benchmark_txn_insert/batchSize=100-4       	    3396	    351671 ns/op	   2843569 objects/sec	  329613 B/op	    6027 allocs/op
Benchmark_txn_delete/batchSize=1-4         	    1572	    766705 ns/op	   1304283 objects/sec	 1270474 B/op	   13976 allocs/op
Benchmark_txn_delete/batchSize=10-4        	    3261	    368897 ns/op	   2710787 objects/sec	  356418 B/op	    5769 allocs/op
Benchmark_txn_delete/batchSize=100-4       	    3651	    328621 ns/op	   3043021 objects/sec	  270754 B/op	    5038 allocs/op
Benchmark_LPM_Lookup-4                     	    9346	    127942 ns/op	   7816027 objects/sec	       0 B/op	       0 allocs/op
Benchmark_LPM_All-4                        	  133652	      9220 ns/op	 108457827 objects/sec	      32 B/op	       1 allocs/op
Benchmark_LPM_Prefix-4                     	  133492	      9013 ns/op	 110950882 objects/sec	      32 B/op	       1 allocs/op
Benchmark_LPM_LowerBound-4                 	  247483	      4879 ns/op	 102482039 objects/sec	     288 B/op	       2 allocs/op
PASS
ok  	github.com/cilium/statedb/lpm	12.014s
goos: linux
goarch: amd64
pkg: github.com/cilium/statedb/part
cpu: AMD EPYC 7763 64-Core Processor                
Benchmark_Set_Singleton_Create-4              	21443167	        55.81 ns/op	      24 B/op	       1 allocs/op
Benchmark_Set_Singleton_Has-4                 	100000000	        11.24 ns/op	       0 B/op	       0 allocs/op
Benchmark_StringMap_Txn_Insert-4              	    7629	    145139 ns/op	   6889959 items/sec	   98250 B/op	    1306 allocs/op
Benchmark_Uint64Map_Random-4                  	    1860	    644888 ns/op	   1550656 items/sec	 1321531 B/op	    6031 allocs/op
Benchmark_Uint64Map_Sequential-4              	    1845	    625026 ns/op	   1599933 items/sec	 1703089 B/op	    5753 allocs/op
Benchmark_Uint64Map_Sequential_Insert-4       	    1954	    590942 ns/op	   1692214 items/sec	 1695084 B/op	    4752 allocs/op
Benchmark_Uint64Map_Sequential_Txn_Insert-4   	    9200	    129107 ns/op	   7745516 items/sec	   90464 B/op	    2031 allocs/op
Benchmark_Uint64Map_Random_Insert-4           	    2097	    571712 ns/op	   1749133 items/sec	 1312906 B/op	    5030 allocs/op
Benchmark_Uint64Map_Random_Txn_Insert-4       	    5906	    192768 ns/op	   5187585 items/sec	  117322 B/op	    2408 allocs/op
Benchmark_Insert_RootOnlyWatch-4              	    9951	    117595 ns/op	   8503789 objects/sec	   75552 B/op	    2036 allocs/op
Benchmark_Insert-4                            	   10000	    116016 ns/op	   8619503 objects/sec	   76096 B/op	    2047 allocs/op
Benchmark_WatchReplace-4                      	   10000	    110417 ns/op	   9056560 objects/sec	   57305 B/op	    1006 allocs/op
Benchmark_Modify-4                            	   10000	    107337 ns/op	   9316450 objects/sec	   58056 B/op	    1007 allocs/op
Benchmark_GetInsert-4                         	    8860	    127332 ns/op	   7853511 objects/sec	   58056 B/op	    1007 allocs/op
Benchmark_Replace-4                           	33092198	        36.42 ns/op	  27460790 objects/sec	       0 B/op	       0 allocs/op
Benchmark_Replace_RootOnlyWatch-4             	32892340	        36.26 ns/op	  27582190 objects/sec	       0 B/op	       0 allocs/op
Benchmark_txn_1-4                             	 7647072	       156.0 ns/op	   6411889 objects/sec	      64 B/op	       3 allocs/op
Benchmark_txn_10-4                            	 9920875	       121.9 ns/op	   8200330 objects/sec	      76 B/op	       2 allocs/op
Benchmark_txn_100-4                           	10942660	       109.1 ns/op	   9169084 objects/sec	      67 B/op	       2 allocs/op
Benchmark_txn_1000-4                          	 9453192	       125.7 ns/op	   7957391 objects/sec	      65 B/op	       2 allocs/op
Benchmark_txn_delete_1-4                      	 5039474	       241.6 ns/op	   4139869 objects/sec	     632 B/op	       3 allocs/op
Benchmark_txn_delete_10-4                     	 9234481	       115.0 ns/op	   8695047 objects/sec	     103 B/op	       1 allocs/op
Benchmark_txn_delete_100-4                    	13352260	        89.45 ns/op	  11179659 objects/sec	      35 B/op	       1 allocs/op
Benchmark_txn_delete_1000-4                   	13991613	        87.70 ns/op	  11402679 objects/sec	      28 B/op	       1 allocs/op
Benchmark_Get-4                               	   59154	     20302 ns/op	  49255669 objects/sec	       0 B/op	       0 allocs/op
Benchmark_GetWatch-4                          	   46032	     26035 ns/op	  38410247 objects/sec	       0 B/op	       0 allocs/op
Benchmark_All-4                               	  157222	      7608 ns/op	 131432699 objects/sec	       0 B/op	       0 allocs/op
Benchmark_Iterator_All-4                      	  141912	      8602 ns/op	 116248197 objects/sec	       0 B/op	       0 allocs/op
Benchmark_Iterator_Next-4                     	  151213	      7940 ns/op	 125942330 objects/sec	     896 B/op	       1 allocs/op
Benchmark_Hashmap_Insert-4                    	   14150	     84426 ns/op	  11844641 objects/sec	   74264 B/op	      20 allocs/op
Benchmark_Hashmap_Get_Uint64-4                	  134943	      8897 ns/op	 112395762 objects/sec	       0 B/op	       0 allocs/op
Benchmark_Hashmap_Get_Bytes-4                 	  107714	     11013 ns/op	  90803348 objects/sec	       0 B/op	       0 allocs/op
Benchmark_Delete_Random-4                     	      43	  25792781 ns/op	   3877054 objects/sec	 2539387 B/op	  102756 allocs/op
Benchmark_find16-4                            	225978972	         5.309 ns/op	       0 B/op	       0 allocs/op
Benchmark_findIndex16-4                       	84585986	        13.74 ns/op	       0 B/op	       0 allocs/op
Benchmark_find64-4                            	295878367	         4.058 ns/op	       0 B/op	       0 allocs/op
Benchmark_findIndex64_hit-4                   	294618501	         4.065 ns/op	       0 B/op	       0 allocs/op
Benchmark_findIndex64_miss-4                  	295391479	         4.058 ns/op	       0 B/op	       0 allocs/op
Benchmark_find4-4                             	426100934	         2.812 ns/op	       0 B/op	       0 allocs/op
Benchmark_findIndex4-4                        	320135000	         3.754 ns/op	       0 B/op	       0 allocs/op
BenchmarkSmallWriteTxn/updates_1-4            	  357777	      3075 ns/op	    3529 B/op	       4 allocs/op
BenchmarkSmallWriteTxn/updates_2-4            	  275232	      4338 ns/op	    4742 B/op	       6 allocs/op
BenchmarkSmallWriteTxn/updates_4-4            	  150288	      6817 ns/op	    7155 B/op	      11 allocs/op
BenchmarkSmallWriteTxn/updates_8-4            	  108494	     11927 ns/op	   11927 B/op	      20 allocs/op
BenchmarkSmallWriteTxn/updates_16-4           	   51804	     27341 ns/op	   21290 B/op	      38 allocs/op
BenchmarkAtomicWatchPointerFirstChannel-4     	17827335	        66.39 ns/op	     120 B/op	       2 allocs/op
BenchmarkAtomicWatchPointerChannel-4          	480801778	         2.497 ns/op	       0 B/op	       0 allocs/op
PASS
ok  	github.com/cilium/statedb/part	57.435s
PASS
ok  	github.com/cilium/statedb/reconciler	0.005s
?   	github.com/cilium/statedb/reconciler/benchmark	[no test files]
?   	github.com/cilium/statedb/reconciler/example	[no test files]
go run ./reconciler/benchmark -quiet
1000000 objects reconciled in 1.65 seconds (batch size 1000)
Throughput 605117.77 objects per second
506MB total allocated, 5010294 in-use objects, 231MB bytes in use

@joamaki
joamaki force-pushed the pr/joamaki/improve-watches branch 2 times, most recently from 11aa908 to 4ffa355 Compare September 7, 2026 08:03
@giorio94
giorio94 self-requested a review September 11, 2026 12:52

@giorio94 giorio94 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Changes look reasonable to me, although I definitely don't have a very deep knowledge of all the statedb internals. These changes are also making the logic a notch more difficult to read and reason about, but I guess that's the tradeoff to try to extract some more performance out of it.

Comment thread part/node.go Outdated
The watch state was still allocated eagerly. This refactors the code
to fully allocate watches lazily. The nodes now contain a 'lazyWatchChannel'
which holds an atomic pointer to 'atomicWatchPointer', so essentially
each node has 'atomic.Pointer[atomic.Pointer[*hchan]]'. The indirection
is required as a logical watch may be shared when the tree forks. If
we'd just store 'atomic.Pointer[*hchan]' we wouldn't be able to avoid
a double-close when the tree forks as the two instances would not be
able to synchronize.

goos: linux
goarch: arm64
pkg: github.com/cilium/statedb/part
              │  before   │              after              │
              │  sec/op   │  sec/op     vs base             │
_Insert          112.1µ      111.1µ            ~ (p=0.280 n=10)
_WatchReplace   108.39µ      88.58µ      -18.27% (p=0.001 n=10)
_GetWatch        17.91µ      19.75µ      +10.28% (p=0.001 n=10)

              │  before   │              after              │
              │   B/op    │   B/op      vs base             │
_Insert         82.28Ki     74.31Ki       -9.68% (p=0.000 n=10)
_WatchReplace   63.81Ki     55.96Ki      -12.30% (p=0.000 n=10)
_GetWatch         0.000       0.000            ~ (p=1.000 n=10)

              │  before   │              after              │
              │ allocs/op │ allocs/op   vs base             │
_Insert          3.067k      2.047k      -33.26% (p=0.000 n=10)
_WatchReplace    2.011k      1.006k      -49.98% (p=0.000 n=10)
_GetWatch         0.000       0.000            ~ (p=1.000 n=10)

Reconciler benchmark (best of 3) before:

1000000 objects reconciled in 1.38 seconds (batch size 1000)
Throughput 727151.51 objects per second
568MB total allocated, 6015247 in-use objects, 239MB bytes in use

After:
1000000 objects reconciled in 1.34 seconds (batch size 1000)
Throughput 748576.38 objects per second
552MB total allocated, 5010271 in-use objects, 231MB bytes in use

An isolated comparison of the pointer wrapper with the direct channel
pointer shows first channel creation improving from 51.44ns to 42.41ns
(-17.54%), 128B to 120B (-6.25%), and 3 to 2 allocations (-33.33%).
Accessing an existing channel changes from 1.999ns to 2.042ns (+2.15%).

A final comparison against an equivalent atomic.Value implementation on Go
1.25.10 used 15 samples of 300ms each pinned to one CPU. atomic.Value was
31.98% slower when creating the first channel, 6.48% slower when loading an
existing channel, 58.14% slower when closing before observation, and 31.54%
slower for channel creation followed by close. Allocation counts were equal,
but atomic.Value used 128B instead of 120B for channel creation and 16B
instead of 8B for close before observation.

AIL:3
Signed-off-by: Jussi Maki <jussi@isovalent.com>
Avoid the variable-length memequal call for the common zero- and one-byte
compressed prefixes. Keep the existing optimized comparison for longer
prefixes.

name                   old time/op  new time/op  delta
Benchmark_Get-6            14.17us      13.42us  -5.24%
Benchmark_GetWatch-6        19.97us      18.48us  -7.46%
Benchmark_GetInsert-6      83.38us      81.43us  -2.33%

Benchmark_Get throughput improves from 70.59M to 74.49M objects/sec
(+5.53%), and Benchmark_GetWatch improves from 50.09M to 54.13M
objects/sec (+8.07%). Allocation counts and bytes are unchanged.

AIL:3
Signed-off-by: Jussi Maki <jussi@isovalent.com>
Use a dedicated search path for Tree.Get and Txn.Get that does not track
the closest watch while traversing the tree. Keep the existing watch-aware
path for GetWatch.

name             old time/op  new time/op   delta
Benchmark_Get-6     16.59us      13.67us   -17.58%

Lookup throughput improves from 60.28M to 73.13M objects/sec (+21.33%).
Both versions perform zero allocations.

AIL:3
Signed-off-by: Jussi Maki <jussi@isovalent.com>
@joamaki
joamaki force-pushed the pr/joamaki/improve-watches branch from 0b30d05 to 20c68e0 Compare September 14, 2026 09:12
@joamaki
joamaki merged commit 2a72ea2 into main Sep 14, 2026
1 check passed
@joamaki
joamaki deleted the pr/joamaki/improve-watches branch September 14, 2026 09:24
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants