Remote debugging plugins were not being synchronized across cluster nodes,
causing "no plugin available nodes found" errors when trying to invoke
plugins from different nodes.
1. **Remote debugging plugins not registered to cluster** - The
`ClusterTunnel` notifier was not being added to ControlPanel
2. **Plugin ID inconsistency** - Remote plugins used different plugin_id
formats during installation vs. querying
3. **Non-idempotent registration** - `RegisterPlugin` failed on reconnection
with "plugin has been registered" error
- **internal/types/models/curd/atomic.go**:
- Unify plugin_id calculation for remote plugins (author/name without version)
- Remove plugin_id from plugin query conditions
- Clear old cache when plugin_id is updated
- **internal/cluster/plugin.go**:
- Make `RegisterPlugin` idempotent by updating existing plugin instead
of returning error
- **internal/core/control_panel/daemon.go**:
- Add cluster field to ControlPanel
- Add SetCluster() method for lazy cluster initialization
- **internal/core/control_panel/server_debugger.go**:
- Register remote debugging plugins to cluster on connection
- Unregister from cluster on disconnection
- **internal/core/plugin_manager/manager.go**:
- Add SetCluster() method to set cluster after initialization
- **internal/server/server.go**:
- Call SetCluster() instead of AddClusterTunnel()
Only remote debugging plugins are synchronized across cluster nodes.
Local plugins run only on the node where they are installed and are
not registered to the cluster.
- Error handling improvements using `errors.Is()` instead of `==`
- Handle 404 for missing plugin assets gracefully
- Handle already-installed debugging plugins gracefully
- Remote debugging plugin can be invoked from any node in the cluster
- Plugin reconnection works without errors
- Cache invalidation works correctly when plugin_id changes
* feat: add retry mechanism for serverless invocation on 502 errors
- Add MAX_SERVERLESS_RETRY_TIMES config (default: 3) for configurable retry attempts
- Implement exponential backoff retry logic (500ms, 1s, 2s, ...) for 502 Bad Gateway errors
- Only retry on 502 status code as it indicates transient AWS Lambda gateway issues
- Add comprehensive test suite with 11 test cases covering all retry scenarios
- Ensure proper HTTP response body cleanup on retries to prevent resource leaks
* refactor: simplify retry logic and error handling in serverless invocation
* fix: max time
* fix: test
* feat(#450): add Redis SSL/TLS configuration support
Add comprehensive SSL/TLS support for Redis connections with configurable certificate verification modes. Introduces new environment variables for SSL configuration including REDIS_USE_SSL, REDIS_SSL_CERT_REQS (supporting CERT_NONE, CERT_OPTIONAL, CERT_REQUIRED), and REDIS_SSL_CA_CERTS for custom CA certificates.
Changes:
- Add Redis SSL configuration options to .env.example
- Implement RedisTLSConfig() method to build tls.Config based on environment settings
- Pass TLS config to both standard Redis and Sentinel mode initializers
- Support custom CA certificate loading and verification modes
- Set minimum TLS version to 1.2 for security
- Minor whitespace cleanup in existing config comments
This enables secure Redis connections in production environments with flexible certificate verification options.
* fix(#450): prevent reference cycle in TLS config and simplify SSL setup
- Capture only RootCAs in VerifyConnection closure to avoid retaining
entire tlsConf and potential reference cycles
- Remove redundant nil checks for tlsConf in Redis client initialization
since tlsConf is guaranteed to be non-nil when useSsl is true
- Update comments to reflect actual behavior and constraints
* fix(#450): improve Redis TLS certificate verification logic for optional certificates
* fix(#450): simplify Redis TLS certificate verification logic for optional and required certificates
* docs(#450): add note for CA certificate file path in Redis SSL configuration
* test(#450): add comprehensive tests for Redis TLS configuration
* fix(#450): enhance Redis SSL configuration documentation and enforce CA cert requirement
* fix(#450): add nil TLS parameter to InitRedisClient calls in tests
Update all InitRedisClient function calls across test files to include the new nil parameter for TLS configuration. This change maintains backward compatibility by explicitly passing nil for TLS settings in non-TLS test scenarios.
* fix(#450): add default TLS configuration for Redis client when no tlsConf is provided
* feat: prioritize `pyproject.toml` when installing plugin dependencies
* refactor: rename variable for clarity in SDK version extraction logic
* refactor: use local_runtime.ConstructPluginRuntime
* refactor: use local_runtime.ConstructPluginRuntime
* refactor: fail when uv not found
* refactor: use more realistic test plugin data
* refactor: don't accept dependencyFileType other than pyprojectTomlFile and requirementsTxtFile
* refactor: use type-safe constant to replace hard-coded string
* refactor: use factory function of LocalPluginRuntime
* use slog instead of log package and format to new log schema
* update the environment name to LOG_OUTPUT_FORMAT
* add the env to .env.example
* fix log reference error
* change the order of milldlewares
* delete unused code
* fix the concurrently session potential race condition
* fix the log format in tests
* update the duplicate code
* refactor: convert log functions to slog structured format
- Change log.Error/Info/Warn/Debug/Panic to accept msg + key-value pairs
- Remove printf-style formatting from log functions
- Update log calls in internal/cluster, internal/db, internal/core/session_manager
- Remove unused 'initialized' variable from log package
- Remaining files will be updated in follow-up commits
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
* refactor: update all log call sites to use slog structured format
Convert all log.Error, log.Info, log.Warn, log.Debug, and log.Panic
calls from printf-style formatting to slog key-value pairs.
Before: log.Error("failed to do something: %s", err.Error())
After: log.Error("failed to do something", "error", err)
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
* refactor: update cmd/ log calls to use slog structured format
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
* feat: implement GnetLogger for structured logging in gnet
* refactor: remove deprecated log visibility functions and related calls
* feat: enhance session management with trace and identity context propagation
* feat: implement serverless transaction handler and writer for plugin runtime
* refactor: rename context field to traceCtx in RealBackwardsInvocation
---------
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
Co-authored-by: Yeuoly <admin@srmxy.cn>
* fix PluginDaemonInternalServerError: no available node, plugin runtime not found
Fix the bug where after a plugin is successfully upgraded, new requests result in the plugin daemon service responding with "no available node, plugin runtime not found".
The root cause is that when the plugin is upgraded successfully, the corresponding value in Redis is not updated. Consequently, when a new request arrives, it reads the outdated plugin version from Redis and fails to locate the correct runtime.
* raise error if failed
* feat: add support for PgBouncer in database initialization
* refactor: consolidate gorm configuration for database connection
* Add MySQL and multi-driver DB integration tests (#535)
* feat: streamline integration tests by using a centralized docker-compose file
* fix: add timeout
* refactor: replace inline service definitions with docker-compose action for integration tests
* fix: update pgbouncer image and environment variable names for consistency
* add locking to prevent simultaneous installations of the same plugin
* ensure proper unlocking of keys in case of errors during installation and upgrade
* handle database not found error in DeletePluginInstallationItemFromTask
* remove enterprise logics
* release lock if runtime already installed
* support setting serverless endpoint by api
* remove global tenant id totally
* scan timeout tasks
* adding comments
* use index in loop
* add log