When replication fails, it starts over attempting to replicate the same amount of data as before.
Replication has improved in AppAssure 5.4 and Rapid Recovery 6 (these results do not exist in builds below 5.4), so that when it fails it starts off at the last 8KB block it was replicating, rather than the last recovery point (assuming the core service hasn't been restarted and a repository check has not been run). When replication restarts, it should start at 0% with a child job of "checking data", wherein it verifies that both cores are in agreement regarding what data has been replicated. This portion of the replication job should move relatively quickly, and will move up the percent completion as it verifies the replicated data, which is tracked in the Staging Area. Once this completes, the child job will change to "transferring", and it will move forward at regular speeds
If you are on a build earlier than 5.4, upgrade to the most recent build of AppAssure or Rapid Recovery.
If the core service was restarted or a repository check was run before replication is resumed, the Staging Area will be cleared, and replication will pick back up at the start of the next recovery point that needs to be replicated
If the above conditions are not applicable, this problem is caused by insufficient space in the Staging Area. The Staging Area is controlled by the deduplication cache, and both are directly proportional to the amount of unique data in the repository
Process Read-Match-Write Algorithm (RMW):
Replication Interruption
Let’s make a simple example of replication between Core1 and Core2 with a 10GB transfer. RMW has completed to step 2 and data is being transferred as described in step 3. Replication is 50% complete when there is an ISP service interruption. What happens?
Keys already sent from step 3 are stored in cache for up to one hour. If communication resumes within one hour:
All SHA Keys are retransmitted for verification, steps 1 & 2
Replication resumes sending data starting at the 51% mark, step 3
If communication does not resume within one hour, the cache is cleared, and Replication must begin anew.
To increase the dedupe cache size: