Repository navigation
agent_persistent by default not persistent #4665
Description
Activity
@klirichek, please prepare a specification for improving this task as you see fit.
Persistent connections are just 'possibility', not 'requirement'. So, they don't need so hard exclusive rules.
Persistent connection, as any other, has attached timeout. When reached, connection is closed despite of persistence. The fact is that this period is longer than for usual (non-persistent) connection. Persistent connections sometimes may not be established. For example, if connections pool is exchaused. In this case usual (non-persistent) connection is in use, and that is no reason to loudly speak about it. Next try, it might be that any other connection was released and returned to the pool, and so, next try it might be available and work as expected. So, that is no hard rules for it.
Also, N of sockets is limited resource, most obviously by N of available ports.Con: behavior change; every distributed table opens persistent pools by default (more idle connections/resources).
That is strange statement, and it is wrong. Distributed table NEVER opens persistent pool 'by default', and NEVER make any default idle connections/resources. Persistent pool is just lazy storage of sockets, and it is initially empty. 'persistent_connections_limit' defines, literally, size limit of such pool. It doesn't allocate the pool, and moreover, it doesn't open any connection.
Generic way distributed table works with remote agent is simple: it just creates the socket, opens connection, uses it, and finally closes.
Connection pool is just a storage. By default it has size limit (that is
persistent_connections_limit), and by default it is empty.
Each remote host (declared as agent) has own connection pool. It is defined by endpoint (host:port). Right nowpersistent_connections_limitis defined per-pool. That is, if you have overall 100 endpoints over all your distributed tables, and connection limit=222, that means, that the instance theoretically can have 22200 persistent connections, distributed over the pools (i.e. 222 connections per each of 100 endpoints). That is no special decision under this model, it is just 'it is made' (i.e., if might be implemented as shared limit per all pools, or as single pool for all - but historically it is as descirbed, i.e. one pool per endpoint, and each pool has own connections limit).'Persistent' way distributed table works with remote agents is this: it rents a socket from persistent pool, opens connection (if necessary), uses it, and finally returns the connection back to the pool.
Pool of persistent socket initially has just limit (say, let's limit be 222, to be clear). That limit illustrates size of internal cyclic buffer (FIFO) of sockets. Renting the socket gives you the socket from the beginning of the buffer; returning will push one to the tail. Initially (when there are no sockets in the pool) it just returns you an empty socket, and count N of the rented. I.e., if you rent, rent, rent - it will give you 222 empty sockets, and finally will say, limit exchaused.
So, if limit is globally set to 0 - the pool will immediately signal 'limit exhaused' to the very first rent request.
Distributed table rents the socket from the pool, and checks it. If it receives already connected socket - it just uses it, and that is the goal the pool exists for. If it receives an empty socket - it creates new connection and establish persistent state. In both case, after use the table will return socket back to the pool. When returning previously empty, the pool grows it's storage by this newcome socket. In case, pool has no more connections to rent (when limit exhaused), distributed table will use plain (non-persistent) connection instead, it will close it after use, and will not return it to the pool (because it wasn't rented from the pool).
Initially pool is empty, and it just return up to connections_limit of new sockets as answer too 'rent' request. So, if we have limit = 222, then for first 222 requests of 'rent socket' pool will just returns '-1', which internally means 'empty socket'. When the limit reached, if will return '-2', which internally means 'no more sockets available'.
If distributed table receives empty socket (-1) from the pool, it open the connection, establish persistent command, then uses it, and finally returns to the pool without closing. If it recieves 'no socket' (-2), it just resets agent state from 'agent_persistent' to simple 'agent' - and then behave as with usual socket without persistence - i.e, creates the socket, opens connection, uses it, and finally closes. It returns nothing to the pool, because it rented nothing.
So, the actual storage of 'connections' is fullfilled only by returned previously rented sockets. If nothing is returned (i.e., you rented 222 connections, but returned nothing) - there is no storage; all is in use.
BTW, agents_persistent is actually useless without connections pool. Because connections need to be stored somewhere. So, it is reasonable to either hardcode, say 1 as the limit, or to calculate N of actual users for any endpoint, and init the pool with calculated N of customers.
Both variants are zero-cost.TEST ( PersistentConnectionsPool, functionality )
illustrates how the pool is working.
Bug Description:
The Problem
agent_persistentonly works ifpersistent_connections_limit> 0, but that settingdefaults to 0. So out of the box, a distributed table declared with
agent_persistentis silently turned into a plain
agent:SHOW CREATE TABLEthen reportsagent, notagent_persistent— the user's settingis dropped from the schema with no hard signal. Looks like the table "changed itself".
This is the core complaint: a user-specified setting is silently mutated.
Decision needed (pick one, or A+B)
A — Change the default. Make
persistent_connections_limitdefault to a small non-zero value.agent_persistentjust works, no surprise.B — Fail loud instead of silently downgrading. When
persistent_connections_limit= 0 andagent_persistentis used, return an ERROR (reject CREATE TABLE / table load) instead ofswapping to
agent.Manticore Search Version:
Latest Dev
Operating System Version:
Ubuntu Latest
Have you tried the latest development version?
None
Internal Checklist:
To be completed by the assignee. Check off tasks that have been completed or are not applicable.
Details