Klever-Go KVM: Throttler slot leak in trie account-data sync causes epoch bootstrap / state sync DoS
The account-data trie syncers leak bounded throttler slots on error paths in syncDataTrie(). Each failed trie sync permanently consumes one slot from
the NumGoRoutinesThrottler, and the slot is never returned unless the sync succeeds or the root hash was already present.
I confirmed this on the current default branch develop at commit 9640d63 (observed on May 20, 2026). I also confirmed the bug with a runtime PoC
using the real timeout path in trieSyncer.StartSyncing(): two timed-out sync attempts are enough to exhaust a throttler with capacity 2.
This affects the epoch bootstrap path because syncUserAccountsState() and syncKappAccountsState() create bounded throttlers and abort bootstrap
immediately if the syncer returns an error. Once enough trie-root sync attempts fail, the syncer cannot make forward progress and bootstrap fails.
data/syncer/userAccountsSyncer.godata/syncer/kappAccountsSyncer.godata/trie/sync.gocore/throttler/numGoRoutinesThrottler.gocore/bootstrap/process.goVerified on:
develop HEAD 9640d63Please check whether the same code is present in supported 1.7.x releases.
High
Both account-data syncers call StartProcessing() before creating / starting the trie syncer, but they only call EndProcessing() on the success path
and on the duplicate-root early return.
userAccountsSyncer.syncDataTrie():
func (u *userAccountsSyncer) syncDataTrie(rootHash []byte, ssh data.SyncStatisticsHandler, ctx context.Context) error {
u.throttler.StartProcessing()
u.syncerMutex.Lock()
if _, ok := u.dataTries[string(rootHash)]; ok {
u.syncerMutex.Unlock()
u.throttler.EndProcessing()
return nil
}
dataTrie, err := trie.NewTrie(...)
if err != nil {
u.syncerMutex.Unlock()
return err
}
trieSyncer, err := trie.NewTrieSyncer(arg)
if err != nil {
u.syncerMutex.Unlock()
return err
}
u.syncerMutex.Unlock()
err = trieSyncer.StartSyncing(rootHash, ctx)
if err != nil {
return err
}
u.throttler.EndProcessing()
return nil
}
The same bug exists in kappAccountsSyncer.syncDataTrie().
After StartProcessing(), the following error paths return without EndProcessing():
NumGoRoutinesThrottler is a strict bounded counter:
func (ngrt *NumGoRoutinesThrottler) CanProcess() bool {
valCounter := atomic.LoadInt32(&ngrt.counter)
return valCounter < ngrt.max
}
func (ngrt *NumGoRoutinesThrottler) StartProcessing() {
atomic.AddInt32(&ngrt.counter, 1)
}
func (ngrt *NumGoRoutinesThrottler) EndProcessing() {
atomic.AddInt32(&ngrt.counter, -1)
}
Once leaked, a slot remains consumed for the lifetime of that throttler instance.
The parent loops in both syncers wait for capacity before starting the next account-data trie sync:
for !u.throttler.CanProcess() {
select {
case <-time.After(timeBetweenRetries):
continue
case <-ctx.Done():
return common.ErrTimeIsOut
}
}
So after enough failures, further roots stop progressing and the sync operation eventually returns time is out.
Epoch bootstrap uses these syncers directly and aborts on any error:
err = e.syncUserAccountsState(e.epochStartMeta.Header.TrieRoot)
if err != nil {
return nil, nil, err
}
err = e.syncKappAccountsState(e.epochStartMeta.Header.KAppsTrieRoot)
if err != nil {
return nil, nil, err
}
The throttlers for these paths are real bounded throttlers created from numConcurrentTrieSyncers.
I verified the bug with the real timeout path, not only with a canceled context.
The PoC below uses:
After the first failed sync, one slot remains leaked. After the second failed sync, the throttler is exhausted.
package syncer
import (
"context"
"testing"
"time"
commonmock "github.com/klever-io/klever-go/common/mock"
corethrottler "github.com/klever-io/klever-go/core/throttler"
"github.com/klever-io/klever-go/data"
"github.com/klever-io/klever-go/data/trie"
triestats "github.com/klever-io/klever-go/data/trie/statistics"
"github.com/stretchr/testify/require"
)
func newBaseSyncerForTimeoutPOC(t *testing.T) *baseAccountsSyncer {
t.Helper()
storageManager, err := trie.NewTrieStorageManagerWithoutPruning(commonmock.NewMemDbMock())
require.NoError(t, err)
return &baseAccountsSyncer{
hasher: commonmock.HasherMock{},
marshalizer: &commonmock.MarshalizerMock{},
trieSyncers: make(map[string]data.TrieSyncer),
dataTries: make(map[string]data.Trie),
trieStorageManager: storageManager,
requestHandler: &commonmock.RequestHandlerStub{},
timeout: time.Second,
cacher: commonmock.NewCacherStub(),
maxTrieLevelInMemory: 5,
name: "timeout-poc",
maxHardCapForMissingNodes: 1,
}
}
func TestPOC_UserAccountsSyncer_LeaksThrottlerSlotOnTrieTimeout(t *testing.T) {
thr, err := corethrottler.NewNumGoRoutinesThrottler(2)
require.NoError(t, err)
s := &userAccountsSyncer{
baseAccountsSyncer: newBaseSyncerForTimeoutPOC(t),
throttler: thr,
}
err = s.syncDataTrie([]byte("missing-root-1"), triestats.NewTrieSyncStatistics(), context.Background())
require.ErrorIs(t, err, trie.ErrTimeIsOut)
require.True(t, thr.CanProcess())
err = s.syncDataTrie([]byte("missing-root-2"), triestats.NewTrieSyncStatistics(), context.Background())
require.ErrorIs(t, err, trie.ErrTimeIsOut)
require.False(t, thr.CanProcess())
}
func TestPOC_KappAccountsSyncer_LeaksThrottlerSlotOnTrieTimeout(t *testing.T) {
thr, err := corethrottler.NewNumGoRoutinesThrottler(2)
require.NoError(t, err)
s := &kappAccountsSyncer{
baseAccountsSyncer: newBaseSyncerForTimeoutPOC(t),
throttler: thr,
}
err = s.syncDataTrie([]byte("missing-root-1"), triestats.NewTrieSyncStatistics(), context.Background())
require.ErrorIs(t, err, trie.ErrTimeIsOut)
require.True(t, thr.CanProcess())
err = s.syncDataTrie([]byte("missing-root-2"), triestats.NewTrieSyncStatistics(), context.Background())
require.ErrorIs(t, err, trie.ErrTimeIsOut)
require.False(t, thr.CanProcess())
}
go test ./data/syncer -run 'TestPOC_(User|Kapp)AccountsSyncer_LeaksThrottlerSlotOnTrieTimeout' -count=1
ok github.com/klever-io/klever-go/data/syncer 4.005s
This confirms the leak with the real timeout path from trieSyncer.StartSyncing().
An attacker who can repeatedly cause trie-node sync failures or timeouts during bootstrap can consume the bounded sync throttler until no capacity remains.
Once enough slots are leaked:
This is a core node availability issue. It affects fresh/restarting nodes and validators that need to bootstrap or resync state.
This is not a theoretical issue:
Release the slot with defer immediately after StartProcessing() and cancel the defer only if ownership is intentionally transferred, which is not the case here.
Example fix pattern:
func (u *userAccountsSyncer) syncDataTrie(rootHash []byte, ssh data.SyncStatisticsHandler, ctx context.Context) error {
u.throttler.StartProcessing()
defer u.throttler.EndProcessing()
u.syncerMutex.Lock()
defer u.syncerMutex.Unlock()
if _, ok := u.dataTries[string(rootHash)]; ok {
return nil
}
dataTrie, err := trie.NewTrie(...)
if err != nil {
return err
}
trieSyncer, err := trie.NewTrieSyncer(arg)
if err != nil {
return err
}
u.trieSyncers[string(rootHash)] = trieSyncer
return trieSyncer.StartSyncing(rootHash, ctx)
}
The same pattern should be applied to:
왜 이 VPI인가 (설명가능 · 실험적)
VPI 산정 기준
| 영향도 | 59.00 |
| 악용 신호(추가 악용신호 없음) | ×1.00 |
| VPI | 59.00 |
VPI 공식 vpi-v1 기준