Skip to content

agent_persistent by default not persistent #4665

Description

@donhardman

Bug Description:

The Problem

agent_persistent only works if persistent_connections_limit > 0, but that setting
defaults to 0. So out of the box, a distributed table declared with agent_persistent
is silently turned into a plain agent:

  • The daemon only logs a warning and falls back to non-persistent.
  • SHOW CREATE TABLE then reports agent, not agent_persistent — the user's setting
    is dropped from the schema with no hard signal. Looks like the table "changed itself".

This is the core complaint: a user-specified setting is silently mutated.

Decision needed (pick one, or A+B)

A — Change the default. Make persistent_connections_limit default to a small non-zero value.

  • Pro: agent_persistent just works, no surprise.
  • Con: behavior change; every distributed table opens persistent pools by default (more idle connections/resources).

B — Fail loud instead of silently downgrading. When persistent_connections_limit = 0 and
agent_persistent is used, return an ERROR (reject CREATE TABLE / table load) instead of
swapping to agent.

  • Pro: no silent schema mutation; user is told exactly what to fix.
  • Con: stricter; configs that relied on the silent fallback now break.

Manticore Search Version:

Latest Dev

Operating System Version:

Ubuntu Latest

Have you tried the latest development version?

None

Internal Checklist:

To be completed by the assignee. Check off tasks that have been completed or are not applicable.

Details
  • Implementation completed
  • Tests developed
  • Documentation updated
  • Documentation reviewed

Activity

  1. sanikolaev commented on Jun 26, 2026

    @sanikolaev
    Collaborator

    @klirichek, please prepare a specification for improving this task as you see fit.

  2. klirichek commented on Jun 30, 2026

    @klirichek
    Contributor

    Persistent connections are just 'possibility', not 'requirement'. So, they don't need so hard exclusive rules.

    Persistent connection, as any other, has attached timeout. When reached, connection is closed despite of persistence. The fact is that this period is longer than for usual (non-persistent) connection. Persistent connections sometimes may not be established. For example, if connections pool is exchaused. In this case usual (non-persistent) connection is in use, and that is no reason to loudly speak about it. Next try, it might be that any other connection was released and returned to the pool, and so, next try it might be available and work as expected. So, that is no hard rules for it.
    Also, N of sockets is limited resource, most obviously by N of available ports.

    Con: behavior change; every distributed table opens persistent pools by default (more idle connections/resources).

    That is strange statement, and it is wrong. Distributed table NEVER opens persistent pool 'by default', and NEVER make any default idle connections/resources. Persistent pool is just lazy storage of sockets, and it is initially empty. 'persistent_connections_limit' defines, literally, size limit of such pool. It doesn't allocate the pool, and moreover, it doesn't open any connection.

    Generic way distributed table works with remote agent is simple: it just creates the socket, opens connection, uses it, and finally closes.

    Connection pool is just a storage. By default it has size limit (that is persistent_connections_limit), and by default it is empty.
    Each remote host (declared as agent) has own connection pool. It is defined by endpoint (host:port). Right now persistent_connections_limit is defined per-pool. That is, if you have overall 100 endpoints over all your distributed tables, and connection limit=222, that means, that the instance theoretically can have 22200 persistent connections, distributed over the pools (i.e. 222 connections per each of 100 endpoints). That is no special decision under this model, it is just 'it is made' (i.e., if might be implemented as shared limit per all pools, or as single pool for all - but historically it is as descirbed, i.e. one pool per endpoint, and each pool has own connections limit).

    'Persistent' way distributed table works with remote agents is this: it rents a socket from persistent pool, opens connection (if necessary), uses it, and finally returns the connection back to the pool.

    Pool of persistent socket initially has just limit (say, let's limit be 222, to be clear). That limit illustrates size of internal cyclic buffer (FIFO) of sockets. Renting the socket gives you the socket from the beginning of the buffer; returning will push one to the tail. Initially (when there are no sockets in the pool) it just returns you an empty socket, and count N of the rented. I.e., if you rent, rent, rent - it will give you 222 empty sockets, and finally will say, limit exchaused.

    So, if limit is globally set to 0 - the pool will immediately signal 'limit exhaused' to the very first rent request.

    Distributed table rents the socket from the pool, and checks it. If it receives already connected socket - it just uses it, and that is the goal the pool exists for. If it receives an empty socket - it creates new connection and establish persistent state. In both case, after use the table will return socket back to the pool. When returning previously empty, the pool grows it's storage by this newcome socket. In case, pool has no more connections to rent (when limit exhaused), distributed table will use plain (non-persistent) connection instead, it will close it after use, and will not return it to the pool (because it wasn't rented from the pool).

    Initially pool is empty, and it just return up to connections_limit of new sockets as answer too 'rent' request. So, if we have limit = 222, then for first 222 requests of 'rent socket' pool will just returns '-1', which internally means 'empty socket'. When the limit reached, if will return '-2', which internally means 'no more sockets available'.

    If distributed table receives empty socket (-1) from the pool, it open the connection, establish persistent command, then uses it, and finally returns to the pool without closing. If it recieves 'no socket' (-2), it just resets agent state from 'agent_persistent' to simple 'agent' - and then behave as with usual socket without persistence - i.e, creates the socket, opens connection, uses it, and finally closes. It returns nothing to the pool, because it rented nothing.

    So, the actual storage of 'connections' is fullfilled only by returned previously rented sockets. If nothing is returned (i.e., you rented 222 connections, but returned nothing) - there is no storage; all is in use.

  3. klirichek commented on Jun 30, 2026

    @klirichek
    Contributor

    BTW, agents_persistent is actually useless without connections pool. Because connections need to be stored somewhere. So, it is reasonable to either hardcode, say 1 as the limit, or to calculate N of actual users for any endpoint, and init the pool with calculated N of customers.
    Both variants are zero-cost.

  4. klirichek commented on Jun 30, 2026

    @klirichek
    Contributor

    TEST ( PersistentConnectionsPool, functionality )
    illustrates how the pool is working.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions