# Configuring Multi-threaded Continuous Monitoring for Sonatype Lifecycle

This guide provides instructions for sizing and enabling multi-threaded Continuous Monitoring (CM) for Sonatype Lifecycle (IQ Server). Multi-threading can significantly reduce CM processing time, helping CM complete within a desired execution window while maintaining system stability and performance.

## Key Takeaways

- CM is single-threaded by default and runs on one IQ Server node in HA.
- Use the total single-threaded CM time (from logs) and divide by 20 hours to compute an initial thread pool size: `-Dinsight.threads.monitor=<N>`.
- Start small (2 threads) and follow CPU guidance (≤ 1/4 of CPUs). Do not exceed 20.
- Set the JVM property on every HA node and restart the server(s) for the change to take effect.

## Overview

The CM job evaluates application policy activity daily (defaulting to 12:00 AM). While CM is single-threaded by default and runs on a single node in High Availability (HA) environments, multi-threading can significantly reduce total processing time. However, it requires careful configuration to avoid excessive CPU utilization or service throttling.

## When to Consider Multi-threading

| Total CM Duration | Recommended Action |
| --- | --- |
| Under 24 hours | No configuration changes required. |
| 24–36 hours | Optimize scanned policies and monitored stages before adjusting threads. |
| Exceeds 36 hours | Proceed with multi-threaded configuration after applying existing optimizations. |

### Scheduling and High-availability (HA) Behavior

#### What Happens When CM Is Already Running

- CM never starts a second instance while another run is active. The scheduler checks for an active CM job and skips scheduling a new one if one is already running.
- This prevents overlapping CM jobs and avoids unnecessary load on the server.

#### How CM Runs in an HA Cluster

- In a high-availability (HA) cluster, only one node runs the scheduled CM job at a time. That node is the active CM node.
- The other nodes continue to serve normal requests but do not run the scheduled CM job while the active node is processing it.

#### JVM Property to Set

Add this JVM argument to the JVM options on every IQ Server node:

```
-Dinsight.threads.monitor=<N>
```

This property controls how many threads the CM worker uses on the node that actually runs the job.

> **Note**  
> The property is read at node startup and cached in memory. Changing the JVM argument requires restarting the node.

#### Why You Must Configure the Property on Every Node

- Each node reads the JVM property when it starts. If a node later becomes the active CM node (for example, after failover), it will use the value it read at startup.
- If a node does not have `-Dinsight.threads.monitor=<N>` set, it falls back to the default of one thread. That can make CM run much slower after failover.
- Configure the property on every node so CM behavior is consistent and predictable across failover and maintenance events.

## Determining Current CM Performance from Logs

Analyze the `clm-server.log` to identify the single-threaded baseline. Search for the following log patterns:

| Logger | Message Pattern | Description |
| --- | --- | --- |
| PolicyMonitor | Starting policy monitoring | Indicates the global start of the CM job. |
| PolicyMonitor | Policy monitoring evaluated for application `<app>` in `<N>` ms | Logs the completion of an individual application scan. |
| PolicyMonitor | Finished policy monitoring applications in `<N>` ms | Provides the authoritative total duration for the job. |

To verify the configured thread pool at startup, look for a line similar to:

```
com.sonatype.insight.brain.policy.evaluator.PolicyMonitor - insight.threads.monitor pool-size: 1
```

## Calculating the Optimal Thread Pool Size

To ensure CM completes within a 20-hour window (leaving a 4-hour operational buffer), calculate the thread pool size using the total single-threaded duration extracted from the logs.

Formula:

```
OptimalThreadPoolSize = ceil( TotalTimeInMs / (20 × 60 × 60 × 1000) )
```

- **TotalTimeInMs** = the single-threaded total time from the log line: _Finished policy monitoring applications in ms_
- **20 × 60 × 60 × 1000** = 72,000,000 ms (20 hours).

Example

```
Log: Finished policy monitoring applications in 437824876 ms
Allowed time (20 hours) = 72,000,000 ms
Optimal pool size = ceil(437,824,876 / 72,000,000) = 7
=> java argument: -Dinsight.threads.monitor=7
```

## Thread Pool Calculation Helper Script

Paste this into Bash/Zsh to extract the `<N>` ms value from your log line and compute the suggested thread pool size.

```
calculate_thread_pool_size() {
  local log_line="$1"
  local buffer_hours=4
  local total_hours=24
  local allowed_hours=$((total_hours - buffer_hours))
  local allowed_time_ms=$((allowed_hours * 60 * 60 * 1000))
  local total_scan_time_ms
  local pool_size

total_scan_time_ms=$(echo "$log_line" | grep -oE '[0-9]+[[:space:]]*ms' | grep -oE '[0-9]+')

if [[ -z "$total_scan_time_ms" ]]; then
    echo "Error: Could not extract scan time from log line."
    return 1
  fi

pool_size=$(( (total_scan_time_ms + allowed_time_ms - 1) / allowed_time_ms ))

echo "Total scan time: ${total_scan_time_ms} ms"
  echo "Allowed scan time: ${allowed_hours} hours (${allowed_time_ms} ms)"
  echo "Optimal thread pool size: $pool_size"
  echo "java argument: -Dinsight.threads.monitor=$pool_size"
}
```

Example usage

```
log="Finished policy monitoring applications in 437824876 ms"
calculate_thread_pool_size "$log"
```

## Thread Pool Configuration Details and Constraints

The thread pool size is controlled by a JVM system property set at server startup:

```
-Dinsight.threads.monitor=<N>
```

| Property | Value / Notes |
| --- | --- |
| Minimum | 1 (default) |
| Maximum | 20 |
| Read | On node startup only (cached in memory). Changing requires restart. |
| HA | Must be specified on each IQ Server node in HA mode. |

## How to Apply

### Standalone / Linux Command Line

```
java -Dinsight.threads.monitor=N -jar nexus-iq-server-<version>.jar server /etc/nexus-iq-server/config.yml
```

Replace N with the computed pool size.

### Kubernetes (Deployment YAML Example)

```
  - name: JAVA_OPTS
    value: "-Xms2g -Xmx4g -Dinsight.threads.monitor=4"
```

Adjust memory and thread values for your environment.

## Deployment Best Practices

- **Incremental Scaling:** Start with 2 threads and monitor system performance before further increases.
- **CPU Utilization:** Limit thread count to approximately 25% of available CPU cores (e.g., 2 threads for an 8-core system).
- **Resource Profile:** Increasing threads is CPU-intensive, not memory-intensive.
- **Upper Limit:** Never exceed 20 threads. Over-provisioning leads to service degradation and throttling.
- **Post-Deployment Monitoring:** Observe CPU load and request latency. Revert changes if service quality drops.

## Real-world Examples

| Applications | Threads | CM Duration (Approx.) |
| --- | --- | --- |
| ~3,000 | 8 | ~90 minutes |
| ~3,000 | 16 | ~45 minutes |
| ~40,000 | 1 (default) | ~52 hours |
| ~40,000 | 8 (recommended) | ~6 hours (estimated) |

## Troubleshooting and Guidance

- If CM still runs > 24 hours: reduce stages monitored and policies scanned before increasing threads further.
- If CPU spikes or service throttling occurs after increasing threads: reduce thread count and/or scale CPU on the server.
- Use the log line "Finished policy monitoring applications in ms" as the single-threaded baseline.
- In HA: always set the JVM property on every node and restart nodes after the change.

## Warnings

- The property is read once on startup - changing it without restarting will not take effect.
- Do not exceed the maximum of 20 threads.
- CM executes on a single node in HA; multi-threading increases CPU load on that node - ensure the node has sufficient CPU.

## Operator Checklist

- Extract the single-threaded CM duration from `clm-server.log`: `Finished policy monitoring applications in <N> ms`
- Compute candidate thread count with the helper script or formula.
- Choose a conservative starting value: 2 threads or the computed value bounded by CPU guidance.
- Set the JVM argument on each node: `-Dinsight.threads.monitor=<N>` and restart each node.
- Monitor CPU, ad-hoc response times, and CM duration; adjust downward if service impact is observed.
