hashirama/uvg266

mirror of https://github.com/ultravideo/uvg266.git synced 2024-11-30 04:34:07 +00:00

Author	SHA1	Message	Date
Arttu Ylä-Outinen	19e051ea40	Reduce intra threshold Reduces intra threshold for --rd=0 from 20 to 8. Threshold of 20 increased BD-Rate too much.	2017-07-25 13:26:38 +03:00
Arttu Ylä-Outinen	e9cf15465e	Fix inter cost in bipred The cost of coding MV ref indices and MV direction was added to bitcost but not inter cost. Fixed by adding the extra bits to inter as well.	2017-07-24 15:24:04 +03:00
Arttu Ylä-Outinen	edbe00763e	Drop extra parameter in kvz_image_calc_sad Drops the parameter max_lcu_below which was always set to -1.	2017-07-24 15:21:19 +03:00
Arttu Ylä-Outinen	ffac29061f	Fix extrapolated inter SATD	2017-07-24 15:11:05 +03:00
Arttu Ylä-Outinen	631ef53d2a	Fix inter cost calculations Inter costs are computed using SAD except when fractional motion estimation or bi-prediction is enabled. This commit changes search_pu_inter_ref to recalculate the cost with SATD. Fixes inter/intra cost comparisons since intra costs are always SATD costs.	2017-07-24 15:11:05 +03:00
Arttu Ylä-Outinen	6ce2fb1238	Add pixel offsets to encoder_state_config_tile_t Adds fields offset_x and offset_y to encoder_state_config_tile_t.	2017-07-24 15:11:05 +03:00
Arttu Ylä-Outinen	2380ba0d41	Reduce copying in kvz_get_coeff_cost Changes function kvz_get_coeff_cost to only copy the CABAC contexts and not the whole encoder state. Other threads could be simultaneously using the other parts of the encoder state. Only copying the CABAC fixes a TSan data race warning.	2017-07-24 12:38:41 +03:00
Arttu Ylä-Outinen	24b462f801	Align coefficients to 8 bytes Adds alignment attribute to lcu_coeff_t. The coefficients are sometimes handled as 64-bit integers containing four coefficients so the arrays should be aligned to 8 bytes. Fixes a UBSan error about misaligned reads.	2017-07-24 12:37:37 +03:00
Arttu Ylä-Outinen	5ddb43c6fe	Fix undefined left shifts in rdo Replaces left shifts by multiplications when the operand may be a negative value. Left shift of a negative value is undefined behavior.	2017-07-24 12:35:10 +03:00
Arttu Ylä-Outinen	d1e64ad62b	Fix undefined left shifts Replaces left shifts by multiplications when the operand may be a negative value. Left shift of a negative value is undefined behavior.	2017-07-20 11:15:30 +03:00
Arttu Ylä-Outinen	07b5fb9caf	Fix out-of-bounds read in encoderstate When calling encoder_state_encode_leaf with POC 0, index -1 of the GOP array would be accessed. Fixed by skipping the code for I-frames.	2017-07-20 11:15:30 +03:00
Arttu Ylä-Outinen	8c4a3473a8	Change --owf=auto and --threads=auto selection Changes OWF selection so that it is chosen based on the maximum number of parallel CTUs. Number of threads is limited to prevent overhead from extra threads.	2017-07-20 09:42:28 +03:00
Arttu Ylä-Outinen	4fc9b743c1	Drop an unnecessary pthread_cond_broadcast Drop pthread_cond_broadcast on threadqueue->cond in function kvz_threadqueue_waitfor. The broadcast caused threads to be woken up more often than necessary.	2017-07-19 11:09:30 +03:00
Arttu Ylä-Outinen	14003c6a30	Disable printing PSNR with --no-psnr	2017-07-19 10:38:37 +03:00
Arttu Ylä-Outinen	e90bde5c62	Clarify PSNR output Adds letters Y, U and V to the PSNR output to make it clearer that the printed values are the luma and chroma PSNR.	2017-07-19 10:33:43 +03:00
Arttu Ylä-Outinen	fdb3480b54	Enable strategies for SAO reconstruction Re-enables strategies for SAO reconstruction. They were disabled in commit `ec9ff42`.	2017-07-11 10:35:18 +03:00
Arttu Ylä-Outinen	333dba3884	Add static to SAO strategies	2017-07-11 10:02:01 +03:00
Miika Metsoila	e8cc2d8f6a	Small fixes	2017-07-07 13:58:19 +03:00
Arttu Ylä-Outinen	67a60a35e3	Fix invalid calls to normalize_lcu_weights Changes encoder_state_init_new_frame to only call normalize_lcu_weights when the weights have been written to the array and rate control is enabled. When rate control is disabled, the weights are not used.	2017-07-07 11:05:31 +03:00
Arttu Ylä-Outinen	563bc26e71	Fix out-of-bounds read in AVX2 SAO AVX2 version of SAO loaded offsets with a 256 bit read even though there are only five 32 bit integers.	2017-07-06 13:04:52 +03:00
Arttu Ylä-Outinen	0850b17f96	Drop get_wpp_limit in search_inter WPP limit for motion vectors is now computed inside fracmv_within_tile.	2017-07-05 13:22:53 +03:00
Arttu Ylä-Outinen	2a85f0f5a4	Move hard-coded MV limits to encoder_control_t Adds field max_inter_ref_lcu to encoder_control_t. It is used to set up inter-LCU dependencies in encoder_state_encode_leaf and restrict motion vectors in fracmv_within_tile.	2017-07-05 13:22:53 +03:00
Arttu Ylä-Outinen	bb5354f7e2	Relax inter-CTU dependencies when SAO is off When using WPP and OWF, the first CTU of a row depends on the last CTU of the row below in the reference frame. This is necessary when SAO is enabled since we currently do SAO for a whole CTU row at a time. When SAO is disabled, however, it is unnecessary to wait for the whole row. Changes CTUs to depend only on the CTU below in the reference frame instead of the whole row when WPP and OWF are enabled and SAO disabled. Gives a significant speedup when running on a machine with many CPU cores.	2017-07-05 13:21:06 +03:00
Arttu Ylä-Outinen	1efa2708b2	Do SAO reconstruction for a single CTU at a time Moves SAO reconstruction into encoder_state_worker_encode_lcu instead of doing it in a separate step for the whole CTU row. Reconstruction of the rightmost 10 pixels and bottommost 10 pixels of a CTU is delayed until the neighboring CTU has been deblocked. Doing SAO for the whole CTU row at a time caused unnecessary inter-CTU dependencies when using WPP and OWF. The first CTU of a row would need to wait until SAO was done for the row below in the previous frame. Moving SAO reconstruction to immediately after deblocking each CTU fixes this problem.	2017-07-04 15:14:31 +03:00
Arttu Ylä-Outinen	ec9ff42077	Rewrite SAO recon to handle arbitrary sized blocks Adds width and height parameters to function kvz_sao_reconstruct and changes it to take coordinates in units of pixels. This will be useful for doing SAO for areas smaller than a whole CTU.	2017-06-30 16:09:18 +03:00
Miika Metsoila	dcd7acf4fd	Fixed crash and incorrect info output	2017-06-27 16:05:15 +03:00
Miika Metsoila	f8b6234fdb	Changes to refence lists to behave more like L0/L1 lists from the specification	2017-06-27 16:05:15 +03:00
Arttu Ylä-Outinen	2c66e0bbd2	Fix warnings about invalid reads in AVX2 ipol AVX2 filter functions read pixels in chunks of 8 or 16 bytes. At the end of the block, the read goes out of the bounds of the pixels array. The extra pixels do not affect the result. Fixes valgrind complaining about the invalid reads by allocating 5 extra pixels in kvz_get_extended_block_avx2	2017-06-22 09:37:55 +03:00
Arttu Ylä-Outinen	4d20e156db	Fix handling intra period not multiple of GOP length With low delay GOP structure, it is possible to use an intra period that is not a multiple of the GOP structure length. Commit `00c9f52` changed encoder_state_init_new_frame to reset POC on intra frames. GOP offset, however, was not reset, resulting in invalid POCs and references for the following frames. This commit changes function kvz_encoder_feed_frame so that GOP offset is correctly reset on intra frames.	2017-06-22 09:29:00 +03:00
Arttu Ylä-Outinen	00c9f52bd4	Fix setting picture type when using GOP Changes encoder_state_init_new_frame to set intra frame pictype to KVZ_NAL_IDR_W_RADL even when using GOP.	2017-06-21 13:21:47 +03:00
Arttu Ylä-Outinen	f54a25f112	Fix crash when immediately closing encoder When closing the encoder, the pictures stored in the input frame buffer are freed by repeatedly calling kvz_encoder_feed_frame. If the encoder was closed immediately after opening it, kvz_encoder_feed_frame would be called with an unprepared encoder state. This would trigger an assert. Fixed by changing kvz_encoder_feed_frame so that it does not require the encoder state to be prepared.	2017-06-15 11:57:46 +03:00
Arttu Ylä-Outinen	b74e0458fd	Set inter transform depth to zero Sets max_transform_hierarchy_depth_inter to 0 in SPS. This saves some bits because split_transform_flag does not need to be coded for inter blocks. When SMP and AMP blocks are enabled the depth is set to 1 instead. Otherwise inter split flag would default to 1 for SMP and AMP blocks, resulting in an unnecessary transform split.	2017-06-08 10:08:20 +03:00
Arttu Ylä-Outinen	8dd01ba5a9	Refactor helper functions in search Combines functions lcu_set_intra_mode and lcu_set_inter_pu to a single function. Removes some duplicated code.	2017-06-06 10:32:09 +03:00
Arttu Ylä-Outinen	1bbecf7584	Refactor work tree copy functions Extracts common code shared by work_tree_copy_up and work_tree_copy_down to a separate function.	2017-06-06 10:32:00 +03:00
Arttu Ylä-Outinen	2b169d5d63	Fix crash in kvazaar_close Changes kvazaar_close to stop all threads before freeing encoder states. Fixes a crash when the encoder is closed before all pictures have been encoded.	2017-06-02 10:05:33 +03:00
Arttu Ylä-Outinen	eb9a05b7ef	Fix memory leak Changes kvazaar_close to free the remaining pictures in the the input frame buffer. Fixes a memory leak when the encoder is closed while there are pictures left in the buffer.	2017-06-01 15:39:35 +03:00
Arttu Ylä-Outinen	8b2483ca1c	Combine intra reconstruction functions Replaces function kvz_intra_recon_lcu_luma and kvz_intra_recon_lcu_chroma in intra.c with function kvz_intra_recon_cu. The new function can handle reconstruction for both luma and chroma. Removes some duplicated code.	2017-05-24 12:07:31 +03:00
Arttu Ylä-Outinen	e67fdb853d	Move intra leaf TB recon to a separate function Moves code for intra leaf transform block reconstruction from functions kvz_intra_recon_lcu_luma and kvz_intra_recon_lcu_chroma to a new function intra_recon_tb_leaf. Removes some duplicated code.	2017-05-24 12:07:31 +03:00
Arttu Ylä-Outinen	13d2fdbd21	Drop unused kvz_videoframe_get_cu functions	2017-05-24 11:15:31 +03:00
Arttu Ylä-Outinen	f5eef7f33c	Use luma pixel coordinates in encode_coding_tree Changes functions encode_intra_coding_unit and encode_coding_tree to take coordinate arguments in units of luma pixels instead of 8 px blocks. This should make the code easier to understand.	2017-05-24 11:15:31 +03:00
Arttu Ylä-Outinen	525a5180ff	Combine intra CU encoding functions Merges functions encode_intra_coding_unit and encode_intra_coding_unit_encry. Removes a lot of duplicated code.	2017-05-24 11:12:40 +03:00
Arttu Ylä-Outinen	610c91b0c5	Use luma pixel coordinates in TU coding functions Changes functions encode_transform_unit and encode_transform_coeff to take coordinate arguments in units of luma pixels instead of 4 px blocks. This should make the code easier to understand.	2017-05-23 15:36:16 +03:00
Arttu Ylä-Outinen	2e8838de6e	Fix crash when crypto compiled in but disabled When kvazaar was built with crypto++ but running without using encryption features, kvazaar attempted to delete an uninitialized crypto handle. Fixed by setting the handle to NULL in kvz_encoder_state_init.	2017-05-23 14:01:48 +03:00
Arttu Ylä-Outinen	2f2c281e8e	Fix a memory leak in crypto A CryptoPP::CFB_Mode<CryptoPP::AES>::Encryption was allocated at the beginning of encoder_state_encode_leaf and was never freed. This commit changes encoder_state_worker_encode_lcu to delete the CFB_Mode. Also moves crypto handle from encoder_state_config_tile_t to encoder_state_t so that it can be safely deleted without affecting other threads in the same tile.	2017-05-23 11:51:25 +03:00
Arttu Ylä-Outinen	22155950c1	Rewrite crypto to conform to kvazaar code style	2017-05-23 11:51:25 +03:00
Arttu Ylä-Outinen	6829865190	Fix inline declaration in intra_mode_encryption Moves the inline declaration of intra_mode_encryption before the type and changes it to use the INLINE macro. Inline declaration after type triggered a warning on GCC.	2017-05-23 11:50:32 +03:00
Arttu Ylä-Outinen	5f8e17d4ba	Eliminate a race condition in threadqueue Fixes the order of acquiring locks for the job and its dependency in kvz_threadqueue_job_dep_add. The dependency is locked before the job that depends on it. This is the same order as in threadqueue_worker. Acquiring the locks in different order in kvz_threadqueue_job_dep_add and threadqueue_worker would sometimes result in a deadlock.	2017-05-18 12:25:53 +03:00
Arttu Ylä-Outinen	4b213477f0	Return best MV from inter early terminate When using --me-early-termination=sensitive, early termination of inter search used to always return the starting point if no tested motion vector was good enough to continue the search. This commit changes early_termination to always return the best motion vector and cost found.	2017-05-18 09:05:14 +03:00
Arttu Ylä-Outinen	382636de55	Fix handling too large QPs Changes kvz_config_validate to output an error if the given QP is out of range and changes kvz_set_picture_lambda_and_qp to clip the QP to the valid range if is too large after applying QP offset from GOP structure.	2017-05-17 12:41:51 +03:00
Arttu Ylä-Outinen	de8b59c681	Drop unused function kvz_coefficients_blit	2017-05-12 16:48:30 +03:00
Arttu Ylä-Outinen	bcfa5a3cd9	Add a comment explaining the coefficient order	2017-05-12 16:46:57 +03:00
Arttu Ylä-Outinen	95775a1645	Change coefficient storage order Changes coefficient storage order to a zig-zag order. Reduces unnecessary copying of coefficients to temporary arrays.	2017-05-12 16:46:57 +03:00
Arttu Ylä-Outinen	9395867a9a	Quantize all colors in a single traversal Changes kvz_quantize_lcu_residual to process all three colors in a single traversal of the TU tree.	2017-05-12 16:42:41 +03:00
Arttu Ylä-Outinen	1e58fd6b16	Split kvz_quantize_lcu_residual Splits kvz_quantize_lcu_residual to two functions that handle the TU tree recursion and quantization of a single TU.	2017-05-12 16:42:41 +03:00
Arttu Ylä-Outinen	cc87e0dcc7	Combine luma and chroma quantization functions Replaces functions kvz_quantize_lcu_luma_residual and kvz_quantize_lcu_chroma_residual in transform.c with function kvz_quantize_lcu_residual. The new function can handle any of the YUV colors. Removes some duplicated code.	2017-05-12 16:42:41 +03:00
Arttu Ylä-Outinen	1357dd0599	Pass coeffs through encoder state Changes the way coefficients are passed from kvz_search_lcu to kvz_encode_coding_tree. Drops fields coeff_y, coeff_u and coeff_v in videoframe_t and instead passes them through field coeff in endoder_state_t.	2017-05-12 16:42:41 +03:00
Eemeli Kallio	2cad3173ec	Reduced amount of modes for search_intra_rdo	2017-05-12 15:56:07 +03:00
Arttu Ylä-Outinen	26adef4492	Merge branch 'erp-aqp'	2017-05-12 15:05:24 +03:00
Eemeli Kallio	55e0e65733	Added INLINE to kvz_get_ic_rate and kvz_get_coded_level in rdo.c	2017-05-12 15:03:30 +03:00
Arttu Ylä-Outinen	ee3d4d0e78	Add adaptive QP for 360 degree video Adds option --erp-aqp for enabling adaptive QP for 360 degree video with equirectangular projection. When projected into a spherical surface, the middle part of the video covers relatively larger area than the top and bottom parts. Enabling --erp-aqp sets up a ROI delta QP array which uses higher QPs for the top and bottom of the video and lower QPs for the middle part.	2017-05-11 12:31:53 +03:00
Arttu Ylä-Outinen	79cb3a2fd3	Permit negative QP deltas in ROI Delta QPs should not be arbitrarily restricted to positive values.	2017-05-11 12:13:47 +03:00
Arttu Ylä-Outinen	edfbd6f122	Add field lcu_dqp_enabled to encoder_control_t Delta QPs for LCUs are enabled when either ROI coding or rate control is enabled. Having a single field is simpler than always checking whether ROI or rate control is enabled.	2017-05-11 12:13:47 +03:00
Arttu Ylä-Outinen	2f2405dfe6	Fix crash when PU depth is limited When video width or height was not a multiple of the smallest CU size, no prediction would be performed at the border CUs. Kvazaar would later crash at an assertion failure when attempting to write the bitstream for the CU. Fixed by permitting inter and intra prediction when the CU split is forced, even if CUs of that size would otherwise be disabled.	2017-04-27 10:35:48 +03:00
Arttu Ylä-Outinen	9130b5107c	Change handling of infinite PSNR in encmain Changes encmain to print 999.99 as PSNR when SSE is zero. This behavior is in line with HM. Previously SSE was set to 99 when it was zero.	2017-04-27 10:35:13 +03:00
Arttu Ylä-Outinen	a9c878b535	Fix crash with WPP when threads are disabled When WPP is enabled, a reference to SAO reconstruction job is copied from the wavefront to the main encoder state. However, when threads are disabled, the job is a null pointer and dereferencing it crashes the encoder. Fixed by adding a null pointer check.	2017-04-24 12:59:57 +03:00
Arttu Ylä-Outinen	2991962033	Add reference counting to threadequeue_job_t Both the thread queue and the encoder states hold pointers to the thread queue jobs. It is possible that a job is removed from the thread queue and freed while the encoder state is still using it. This commit adds reference counting to threadqueue_job_t in order to fix the problem. Fixes #161.	2017-04-12 16:13:52 +03:00
Arttu Ylä-Outinen	bd8adff43a	Drop unused defines in threads.h	2017-04-12 03:41:07 -07:00
Arttu Ylä-Outinen	7ab0a7aff2	Fix semaphores on Mac POSIX semaphores are deprecated on Mac. This commit replaces POSIX semaphores by Grand Central Dispatch semaphores when building on Mac.	2017-04-12 03:41:02 -07:00
Arttu Ylä-Outinen	26693e1402	Fix reliance on undefined behaviour in encmain Pthread mutexes were used for synchronization in encmain by locking and unlocking them from different threads. However, according to the POSIX standard, unlocking a mutex from a different thread is undefined behaviour. This commit replaces the mutexes by semaphores which can be used from different threads.	2017-04-12 03:23:58 -07:00
Ari Lemmetti	47a9f0de04	Modify and use FILL_ARRAY macro to prevent warning on GCC 7 Following warning was given and is false positive error: 'memset' used with length equal to number of elements without multiplication by element size [-Werror=memset-elt-size]	2017-04-11 14:04:25 +03:00
Eemeli Kallio	f7e01b8ba1	Fixed error on rd=3	2017-04-05 13:27:14 +03:00
Eemeli Kallio	9f605152ae	Changed intra to use best rough cost when using inter and rd=2	2017-04-05 13:01:32 +03:00
Ari Lemmetti	33ce101ab5	Revert "Use sizeof(uint32_t) to avoid warning in GCC7." Did not fix the problem. This reverts commit `e3c3e74926`.	2017-04-03 20:21:33 +03:00
Ari Lemmetti	e3c3e74926	Use sizeof(uint32_t) to avoid warning in GCC7. error: 'memset' used with length equal to number of elements without multiplication by element size [-Werror=memset-elt-size]	2017-04-03 19:16:09 +03:00
Arttu Ylä-Outinen	df359b8f95	Fix indentation in encode_coding_tree.c Fixes indentation of a for loop that was causing a misleading indentation warning on GCC. Fixes #163.	2017-03-08 22:56:28 +09:00
Pierre-Loup Cabarat	2b8ce5e47c	Add intra prediction modes encryption	2017-03-06 17:27:39 +01:00
Arttu Ylä-Outinen	aae141f2d3	Fix order of frames with --debug When the decoding and presentation orders of pictures are different (with GOP), the frames in YUV debug output would be in the decoding order. This commit changes the kvazaar command line program to store the reconstructed pictures in a buffer so that they can be output in the presentation order. Fixes #101.	2017-02-28 14:09:24 +09:00
Arttu Ylä-Outinen	094b39e7fc	Refactor inter MV/merge candidate selection Adds struct merge_candidates_t for holding the spatial and temporal merge candidates. Changes functions with separate parameters for each candidate to use the struct instead.	2017-02-22 15:56:36 +09:00
Arttu Ylä-Outinen	3409748a8f	Refactor inter MVP candidate selection Adds helper function add_mvp_candidate.	2017-02-22 15:56:27 +09:00
Arttu Ylä-Outinen	ef6503c728	Refactor inter merge candidate selection Adds helper function add_merge_candidate and replaces macro CHECK_DUPLICATE with function is_duplicate_candidate.	2017-02-22 02:50:52 +09:00
Arttu Ylä-Outinen	f12e09bc40	Refactor inter TMVP selection Adds helper function add_temporal_candidate to inter.c.	2017-02-22 02:08:10 +09:00
Arttu Ylä-Outinen	4f88066740	Refactor MV and merge candidate selection Replaces macros APPLY_MV_SCALING and CALCULATE_SCALE with helper functions.	2017-02-22 01:14:16 +09:00
Arttu Ylä-Outinen	db08041d9a	Refactor inter TMVP selection Merges three if-clauses to remove two levels of indentation.	2017-02-21 23:56:01 +09:00
Marko Viitanen	85e2a40da3	Clip scaled motion vectors, scale and td/tb values to appropriate limits Fixes #158.	2017-02-20 15:40:20 +02:00
Ari Koivula	7369f25f64	Bump version to 1.1.0	2017-02-16 20:52:05 +02:00
Ari Lemmetti	b021d2244e	Reduce more unnecessary initializations.	2017-02-16 17:25:26 +02:00
Ari Lemmetti	acd12cba1e	Remove unnecessary memory initialization to zero Values in interval [last_scanpos, 0] are overwritten in following for loop, except for the sig_coeff_inc value.	2017-02-16 16:48:48 +02:00
Ari Koivula	7ff33e1bf2	Fix default reference picture count The default was 3, instead of the intended 1 of the medium preset.	2017-02-13 17:34:28 +02:00
Marko Viitanen	4251607c04	Fix a bug in TMVP reference POC list	2017-02-13 15:19:24 +02:00
Marko Viitanen	4270d451e6	Fixed some errors after rebase	2017-02-13 15:19:24 +02:00
Marko Viitanen	95effb00d0	Disable TMVP in frames with zero L0 references	2017-02-13 15:19:24 +02:00
Marko Viitanen	b4de1878be	Fixed TMVP scaling and candidate selection for B-frames	2017-02-13 15:19:23 +02:00
Marko Viitanen	23be633ad7	Added TMVP merge candidate scaling for L0	2017-02-13 15:19:23 +02:00
Marko Viitanen	e6aa1b9b9a	Renamed get_mv_cand_from_spatial() to get_mv_cand_from_candidates()	2017-02-13 15:19:23 +02:00
Marko Viitanen	1124bb5fd0	Cleaned up TMVP, mv candidate selection working, merge candidate selection not	2017-02-13 15:19:23 +02:00
Marko Viitanen	d65d2ec88d	WIP: add list of POCs used in the image when pushing to reference	2017-02-13 15:19:22 +02:00
Marko Viitanen	6a25cd3248	WIP: work on tmvp on inter	2017-02-13 15:19:22 +02:00
Marko Viitanen	e538a94eda	Enable TMVP with B-frames	2017-02-13 15:19:22 +02:00
Arttu Ylä-Outinen	363b8b49a2	Fix integer overflows with large resolutions Limits video size so that the number of luma and chroma pixels can be stored in an int. Fixes some integer overflows that resulted in segmentation faults.	2017-02-12 11:40:13 +09:00
Arttu Ylä-Outinen	a5a925fc28	Replace timed waits by normal waits in threadqueue Replaces calls to pthread_cond_timedwait with pthread_cond_wait in threadqueue.c. Simplifies code, as there should be no need for the timeout.	2017-02-11 15:42:03 +09:00
Arttu Ylä-Outinen	fd057498fc	Simplify kvz_config_alloc	2017-02-11 15:42:03 +09:00
Arttu Ylä-Outinen	7f7844caad	Fix finalizing uninitialized encoder states Finalization functions for frame and tile encoder states accessed the frame and tile fields of the encoder state even though they might be NULL. This is the case when the initialization of an encoder state fails. Fixed by adding NULL checks.	2017-02-09 14:05:28 +09:00
Arttu Ylä-Outinen	51786eda67	Drop redundant fields in encoder_control_t Some of the fields in encoder_control_t were simply copies of the corresponding fields in kvz_config. This commit drops the copied fields in favor of using the fields in encoder_control_t.cfg directly.	2017-02-09 14:05:28 +09:00
Arttu Ylä-Outinen	6a178dee96	Fix leaking memory when --cqmfile given many times Any previously allocated CQM file name was not freed when allocating memory for the new file name.	2017-02-09 14:05:28 +09:00
Arttu Ylä-Outinen	63a567ad8a	Fix leaking memory when --roi given many times Any previously allocated delta QP array was not freed when allocating a new array.	2017-02-09 14:05:21 +09:00
Arttu Ylä-Outinen	bfd89136a4	Fix ROI delta QP array not getting freed	2017-02-09 13:23:55 +09:00
Arttu Ylä-Outinen	e78a8dfcf5	Copy the kvz_config passed to encoder_open The kvz_config struct is created by the user but kvazaar keeps a pointer to it. It is easy to break things by modifying the configuration outside kvazaar. In addition, kvazaar modifies the struct even though it is has a const modifier. This commit changes the field cfg in encoder_control_t to be a copy of the kvz_config struct instead of a pointer, removing modifications to the const struct and allowing users to do whatever they want with it after opening the encoder.	2017-02-09 13:23:54 +09:00
Ari Koivula	b8e3513a23	Fix crash with sub-LCU frame sizes and WPP The end of slice was being calculated incorrectly, which led to no tile being created inside the slice, which led to an assert triggering. This fixes the wrong end of slice calculation, but also disallows wavefront rows from being created, if there would be only one. The wavefront initialization code assumes there are always more than one row, so the inter-frame dependency doesn't get added properly. Fixes #153.	2017-02-08 21:41:30 +02:00
Ari Koivula	d893474bab	Fix encoder getting stuck on OS-X Main thread was stuck looping on pthread_cond_timedwait because the abs time given on OS-X had already passed and the wait returned immediately without releasing the mutex to allow worker threads to proceed. Fix was to use the gettimeofday, which returns real time instead of monotonic, which is what pthread_cond_timedwait wants.	2017-02-02 17:27:46 +02:00
Ari Koivula	4ceda1908b	Fix OS-X compiler warning rdo.c:475:25: warning: absolute value function 'abs' given an argument of type 'int64_t' (aka 'long long') but has parameter of type 'int' which may cause truncation of value [-Wabsolute-value] current.cost = -abs(quant_cost_in_bits) + (bits << PRECISION_INC); ^ rdo.c:475:25: note: use function 'llabs' instead current.cost = -abs(quant_cost_in_bits) + (bits << PRECISION_INC);	2017-02-01 18:09:17 +02:00
Ari Koivula	c7d536bbcd	Fix OS-X compiler warning cfg.c:1024:74: warning: format specifies type 'size_t' (aka 'unsigned long') but the argument has type 'unsigned long long' [-Wformat] fprintf(stderr, "Too large ROI size: %llu (maximum %zu).\n", size, SIZE_MAX);	2017-02-01 18:09:04 +02:00
Ari Koivula	4467506ef1	Add missing kvz_ prefix	2017-01-31 18:38:02 +02:00
Ari Koivula	ed3bd898fd	Remove Exp-Golomb lookup table This table takes 256kB and isn't used very much. Au revoir!	2017-01-31 18:31:05 +02:00
Ari Koivula	5513744d24	Merge branch 'slices'	2017-01-31 16:14:30 +02:00
Ari Koivula	52904d3e9f	Add --slices=tiles and --slices=wpp This encapsulates tiles or WPP rows into their own slices, making it possible to send them as soon as they are done, instead of waiting for the other substreams to finish and coding the substream offsets in the slice header.	2017-01-31 15:44:23 +02:00
Ari Koivula	0d4d0e869c	Add support for independent slices Not used yet, but they work.	2017-01-31 15:11:50 +02:00
Ari Koivula	46ae382498	Fix bugs with slice header These fixes allow more than one slice to be used to code a picture. - Use correct number of bits to code the slice segment address. - Don't offset_len_minus1 for slices without substreams.	2017-01-31 14:01:59 +02:00
Ari Koivula	f1fc0de2bf	Write slice headers to the parent stream Appending to the child stream doesn't work is the child is a leaf slice state. Simplifies flow by removing distinction between tile and slice. Now that slice headers are written in the parent stream, there is zero difference between tiles and slices from bitstream point of view.	2017-01-31 13:55:05 +02:00
Ari Koivula	04cd875b2c	Move substream finalization to LCU coding job Having some of the termination bits in the LCU coding and some in the substream finalization was needlessly confusing. Doing substream finalization directly after LCU coding makes it easy to verify that the finalization is done correctly. Removes one job per WPP row from the job queue. Removes kvz_cabac_flush, because I don't like bits being put into the bitstream implicitly. Better to have it all in the open.	2017-01-31 13:01:57 +02:00
Ari Koivula	ead490b7b7	Write a new slice NAL for every slice	2017-01-31 12:36:18 +02:00
Ari Koivula	cd496bf50b	Move first_nal_in_au to encoder_state->frame Needed for writing NALs from encoder_state_write_bitstream_children	2017-01-31 12:28:28 +02:00
Arttu Ylä-Outinen	1e6463c08b	Fix inter bipred search When the number of merge candidates was five, biprediction search would read past the bounds of the priority list arrays. Fixed to limit the search to the first four candidates.	2017-01-31 18:23:12 +09:00
Ari Lemmetti	2c069a3e5f	Prevent unnecessary cu search Prevent further analysis as soon as it is known that splitting can not improve cost	2017-01-30 16:21:41 +02:00
Arttu Ylä-Outinen	9b889c3fab	Fix reading ROI files - Checks the return value of fopen when opening the ROI file. Fixes a segfault when the file cannot be opened. - Check that the width and height are positive. Fixes reading past the end of the delta QP array in kvz_set_lcu_lambda_and_qp. - Check for overflow in width * height. Fixes an overflow resulting in a segfault. - Properly check that fscanf succeeds. Fixes silently accepting ROI files that are too short. - Properly close the FILE pointer.	2017-01-29 18:57:27 +09:00
Arttu Ylä-Outinen	46c9a483c3	Fix inter search for small SMP and AMP blocks The function search_pu_inter_ref incorrectly rounded the coordinates of the block to down to a multiple 8 pixels. Small SMP and AMP blocks may start at coordinates that are not multiples of 8. Fixed by removing the rounding. Fixes a failing assert when --mv-constraint is used with --smp or --amp.	2017-01-29 13:34:50 +09:00
Arttu Ylä-Outinen	fb10b56b82	Fix checking if a low delay GOP structure is used Stops assuming that having cfg->gop_lowdelay set means that GOP structure is used since it is possible that cfg->gop_lowdelay is true but cfg->gop_len is zero. Adds checks for cfg->gop_len where needed. Fixes a possible division by zero in kvz_encoder_feed_frame.	2017-01-28 21:56:00 +09:00
Arttu Ylä-Outinen	4f56b04239	Drop an unnecessary conditional Drop a conditional for depth > MAX_DEPTH in search_cu. The depth cannot be greater than MAX_DEPTH (== 3) since an earlier if-clause checks that it is less than MAX_PU_DEPTH (== 4).	2017-01-28 21:35:27 +09:00
Ari Koivula	937a764987	Fix bug in --mv-constraint Subpixel motion estimation return 0-vector when no subpixel vector is within the constraint. Fix is to not call subpixel motion estimation when the integer vector is not within the constraint.	2017-01-26 09:55:57 +02:00
Ari Koivula	4a0121ac42	Add --roi parameter Adds region of interest coding capability. Works by reading a file of delta QP values which will then be applied to each frame at LCU level.	2017-01-26 09:14:14 +02:00
Ari Koivula	6f61836989	Refactor kvz_rdoq_sign_hiding Rename and reorder everything to make more sense. - Moved input tables into their own struct and renamed them to what they actually represent. - Renamed pretty much every variable to comform to our style and to make sense. - Removed the lastCG stuff, as the function already gets passed the last coeff anyway. (it was named width, what the hell?)	2017-01-19 23:58:17 +02:00
Ari Koivula	a85390d0ac	Clean up code using the fixed point frac bit tables This is to prepare for changing the code using the floating point table to use the fixed point table instead. This also allows reducing the size of the fractional part, which was useful for finding every place where the the fixed point presentation is relied upon.	2017-01-19 20:20:51 +02:00
Ari Koivula	24a69c7467	Refactor luma deblocking Changes luma deblocking to use gather and scatter instead of reading to and writing from here and there in memory. Should make them faster and easier to vectorize, or at least cleaner. Splits strong and weak luma deblocking to two functions, as they have almost nothing in common.	2017-01-17 22:13:39 +02:00
Ari Koivula	4cb2fca924	Refactor deblock decision	2017-01-17 19:34:17 +02:00
Arttu Ylä-Outinen	05794c3548	Add missing static to function lambda_to_qp	2017-01-11 15:53:55 +09:00
Arttu Ylä-Outinen	ee518e8ac4	Take header bits into account in rate control	2017-01-11 15:53:55 +09:00
Arttu Ylä-Outinen	c219d3cd94	Fix deblock when CU QP delta is enabled Fixes deblock functions so that they use the correct QP for the filtered edge. Adds field qp to cu_info_t.	2017-01-11 15:53:22 +09:00
Arttu Ylä-Outinen	82a98180e4	Clip LCU lambda to reduce quality fluctuation Limits lambdas for each LCU based on the computed lambda from the previous frame and the frame-level lambda.	2017-01-09 01:24:23 +09:00
Arttu Ylä-Outinen	93172fd251	Use separate alpha, beta and lambda for each LCU Changes rate control to use the alpha and beta values stored in lcu_stats_t instead of the frame-level values when selecting lambda and QP for an LCU.	2017-01-09 01:24:23 +09:00
Arttu Ylä-Outinen	3af4e9cc8a	Allocate bits separately for each LCU Bits are allocated based on the costs of the LCUs in the previous completely coded frame. Breaks deblock when rate control is used.	2017-01-09 01:24:23 +09:00
Arttu Ylä-Outinen	ff5e5ec6d4	Record info about coded LCUs Adds field lcu_stats to encoder_state_config_frame_t. The following data is recorded for each LCU: - number of bits - squared cost - used lambda value - alpha parameter used for rate control - beta parameter used for rate control	2017-01-09 01:24:23 +09:00
Arttu Ylä-Outinen	2a4243acbe	Refactor rate control Moves all code related to setting QP and lambda values to rate_control module.	2017-01-09 01:24:23 +09:00
Arttu Ylä-Outinen	71633889ce	Enable CU QP delta when using rate control When rate control is enabled, enable cu_qp_delta_enabled_flag in PPS with diff_cu_qp_delta_depth set to 0. Also adds code for writing the QP deltas and a new cabac context.	2017-01-09 01:24:23 +09:00
Arttu Ylä-Outinen	640ff94ecd	Use separate lambda and QP for each LCU Adds fields lambda, lambda_sqrt and qp to encoder_state_t. Drops field cur_lambda_cost_sqrt from encoder_state_config_frame_t and renames cur_lambda_cost to lambda.	2017-01-09 01:24:23 +09:00
Arttu Ylä-Outinen	435c387357	Refactor rate control - Defines MIN_LAMBDA and MAX_LAMBDA constants. - Moves resetting state->frame->cur_gop_bits_coded to rate_control.c. - Changes gop_allocate_bits to return the number of bits allocated like pic_allocate_bits does.	2017-01-09 01:24:23 +09:00
Arttu Ylä-Outinen	6c4f2d196a	Move fields from encoder_state_t to frame Moves fields prepared and frame_done from encoder_state_t to encoder_state_config_frame_t.	2017-01-09 01:24:23 +09:00
Arttu Ylä-Outinen	97863cdaa2	Fail encoder init when CQM file cannot be opened	2017-01-08 19:17:43 +09:00
Arttu Ylä-Outinen	db5e750c7f	Fix --threads=auto When --threads=auto was given on the command line, cfg->threads was actually set to zero, disabling threads altogether. Fixed to set cfg->threads to -1, so that the number of threads is chosen automatically.	2017-01-08 17:58:22 +09:00
Ari Koivula	a9e45efcfc	Add a fast lane for byte-aligned bitstream writes The CABAC engine only writes to the bitstream when it has a full byte. These writes are also always byte-aligned, so there is no need to even check for stream alignment. Speedup was around 3% with ultrafast and low QP.	2016-12-23 17:01:44 +02:00
Jaakko Laitinen	deb63f735f	Fix gop disabling	2016-12-20 14:25:13 +02:00
Ari Lemmetti	70a52f0e48	10-bit: add missing bit depth adjustment to ssd	2016-11-17 19:28:04 +02:00
Ari Koivula	fa078102f1	Fix 32bit compilation Got a warning about implicit cast from uint64_t to void*.	2016-11-17 17:53:57 +02:00
Ari Koivula	5ceec06bd3	Merge pull request #148 from Venti-/crypto Crypto	2016-11-16 21:33:55 +02:00
Ari Lemmetti	c31207ea7d	Optimize intra reference building -Add function with reduced logic for the most common case	2016-11-16 18:28:42 +02:00
Ari Koivula	24f2a23ef8	Remove unnecessary crypto state The frame does not need it's own crypto state, since it always has at least one sub tile.	2016-11-16 13:58:41 +02:00
Ari Koivula	8951e34fd2	Change crypto.h stubs to print instead of assert	2016-11-16 13:58:41 +02:00
Wassim Hamidouche	ea82c38906	correct memory allocation	2016-11-16 12:35:28 +02:00
Wassim Hamidouche	da3e2d1d07	resolve parallel encryption	2016-11-16 12:35:28 +02:00
Ari Koivula	b8a618e666	Fix problems with >8 bit input Enforce bit depth promised by --input-bitdepth to avoid crashes when larger values are provided. Do endianess byte swap for all bytes when the buffer gets extended to multiple of 8 pixels, and not just the number of input pixels. Don't swap bytes on a little-endian system.	2016-11-13 19:58:54 +02:00
Ari Koivula	2c005cda25	Fix bug with sub-pixel motion estimation in tiles The width of the tile was being used to index the frame pixel buffer instead of the width of the buffer.	2016-11-07 15:53:52 +02:00
Ari Koivula	78a28e0338	Reformat --help message - Reduce indentation to 6 spaces - Word wrap everything to under 80 characters - Remove defaults from options covered by presets - Add a dash in front of argument descriptions - Add --(no-) to names of parameters that accept it and remove mention of enabling or disabling - Add executable and scripts as a dependancy to make docs	2016-11-04 15:40:28 +02:00
Ari Koivula	d18de19d8a	Fix DTS and PTS not being passed on through lib API Fixes "cur_dts is invalid" warning from FFmpeg.	2016-10-28 19:05:47 +03:00
Ari Koivula	0c41c2ebd6	Make CLI set PTS for each input picture This value is not represented in the HEVC bitstream, which is why it was not set previously. FFmpeg sets and needs it however, so make the CLI set it as well to make sure we handle it correctly.	2016-10-28 19:03:03 +03:00
Ari Koivula	5bf745460d	Re-categorize options in the help message - Move VUI stuff to the bottom - Merge Parallel processing, WPP, Tiles and slices - Add more categories for the other options	2016-10-27 03:26:15 +03:00
Ari Koivula	cb6672b452	Disable WPP when Tiles are enabled Closes #142.	2016-10-27 02:07:10 +03:00
darealshinji	488d042e5f	Bump KVZ_VERSION	2016-10-25 12:32:13 +02:00
Ari Lemmetti	29153ed503	Remove unused variable	2016-10-21 17:28:42 +03:00
Ari Lemmetti	778e46dfd8	Add AVX2 version of SSD	2016-10-21 15:07:53 +03:00
Ari Lemmetti	6f5d7c9e06	Move SSD to strategies	2016-10-21 15:07:23 +03:00
Ari Lemmetti	89b941eab4	Fix typo	2016-10-21 15:07:02 +03:00
Alexis Ballier	1dcc993743	Include i386 & i486 for compiling intel asm. x86_64-pc-linux-gnu-gcc -m32 that I use for building 32bits libraries on amd64 defines only __i386__.	2016-10-14 18:07:37 +02:00
Arttu Ylä-Outinen	5fb7afe8c4	Add --implicit-rdpcm command line parameter. Makes it possible to use lossless coding without implicit residual DPCM.	2016-10-03 20:01:55 +09:00
Arttu Ylä-Outinen	5affc0f527	Use implicit RDPCM in lossless mode. Sets implicit RDPCM flag in SPS when lossy coding is disabled and applies DPCM to intra residual when prediction mode is horizontal or vertical.	2016-10-03 19:31:38 +09:00
Ari Koivula	016dbe0894	Further refine presets The rd-complexity of slow presets is better with a less agressive GOP. Adding the GOP as part of the preset improved BDRate enough, that it didn't make sense anymore to have a veryslow target the best BDRate. Instead, push that responsibility to placebo by making it a little bit faster.	2016-09-29 17:35:12 +03:00
Ari Koivula	31c5ff0f16	Add cross-platform core number detection Well, turns out pthread_num_processors_np isn't standard so we need to do this crap. Threw in hyper threading detection as a bonus.	2016-09-29 00:03:21 +03:00
Ari Koivula	8c7351eac8	Fix lp-gop with depth 1 GOPs with depth 1 had the same structure as those with depth 2: g4d3t1 = 3 2 3 1 g4d2t1 = 2 2 2 1 g4d1t1 = 2 2 2 1 It now results in the correct: g4d1t1 = 1 1 1 1	2016-09-29 00:03:21 +03:00
Ari Koivula	a395aeaac9	Set default settings to those of --preset=medium	2016-09-29 00:03:21 +03:00
Ari Koivula	4388fe0d30	Set presets to ratedistortion-complexity optimized versions	2016-09-29 00:03:20 +03:00
Ari Koivula	facb1e16df	Use -p64 -q22 and --gop=lp-g4d3t1 by default Coding inter without GOP of any kind really isn't a very sensible default. Defaulting to B-GOP of some kind would be more better, but lp-gop is more robust for now.	2016-09-29 00:03:20 +03:00
Ari Koivula	d7391a9593	Improve default for number of parallel frames	2016-09-29 00:03:20 +03:00
Ari Koivula	19d423ab29	Use all available cores by default	2016-09-29 00:03:20 +03:00
Ari Koivula	3f138f087a	Allow non-gop-length --period for lp-gop	2016-09-29 00:03:19 +03:00
Ari Koivula	16790c9f15	Remove number of references from --gop=lp syntax The number of references should be part of the presets, so gop should be defined separately.	2016-09-29 00:03:19 +03:00
Ari Koivula	cbfa824d1a	Merge branch 'simd'	2016-09-27 20:49:45 +03:00
Ari Koivula	14a7bcba25	Use a faster function for clipped inter SAD Use the vectorized general SSE41 inter SAD in AVX reg_sad for shapes for which we don't have AVX versions yet. Also improves speed of --smp and --amp a lot. Got a 1.25x speedup for: --preset=ultrafast -q 27 --gop=lp-g4d3r3t1 --me-early-termination=on --rd=1 --pu-depth-inter=1-3 --smp --amp * Suite speed_tests: -PASS inter_sad: 0.898M x reg_sad(64x63):x86_asm_avx (1000 ticks, 1.000 sec) +PASS inter_sad: 2.503M x reg_sad(64x63):x86_asm_avx (1000 ticks, 1.000 sec) -PASS inter_sad: 115.054M x reg_sad(1x1):x86_asm_avx (1000 ticks, 1.000 sec) +PASS inter_sad: 133.577M x reg_sad(1x1):x86_asm_avx (1000 ticks, 1.000 sec)	2016-09-27 20:48:30 +03:00
Arttu Ylä-Outinen	4313e56c2d	Add --no-rdoq-skip command line switch	2016-09-11 17:40:16 +09:00
Ari Koivula	a7a33b08ec	Remove --slice-addresses from usage message And give a warning if it's used. Slices will have to be implemented at some point, but they aren't yet so let's not advertize them.	2016-09-10 21:06:00 +03:00
Eemeli Kallio	f41e428e5f	Removed kvz_skip_unnecessary_rdoq and reworked --rdoq-skip to skip 4x4 blocks when it is on.	2016-09-09 10:26:07 +03:00
Eemeli Kallio	ed9c0b0416	RDOQ reworked in rdo.c. rdoq_signhide now skips coeffs that are after best_last_idx.	2016-09-09 10:16:51 +03:00
Ari Koivula	02cd17b427	Add faster AVX inter SAD for 32x32 and 64x64 Add implementations for these functions that process the image line by line instead of using the 16x16 function to process block by block. The 32x32 is around 30% faster, and 64x64 is around 15% faster, on Haswell. PASS inter_sad: 28.744M x reg_sad(32x32):x86_asm_avx (1014 ticks, 1.014 sec) PASS inter_sad: 7.882M x reg_sad(64x64):x86_asm_avx (1014 ticks, 1.014 sec) to PASS inter_sad: 37.828M x reg_sad(32x32):x86_asm_avx (1014 ticks, 1.014 sec) PASS inter_sad: 9.081M x reg_sad(64x64):x86_asm_avx (1014 ticks, 1.014 sec)	2016-09-01 21:36:39 +03:00
Ari Koivula	d0512d25c6	Use fixed point in get_mvd_coding_cost	2016-08-30 21:37:12 +03:00
Ari Koivula	ec7507a935	Further optimize get_ep_ex_golomb_bitcost Unrolled 16-bit log2 calculation.	2016-08-30 21:37:01 +03:00
Ari Koivula	a4ba794587	Optimize get_ep_ex_golomb_bitcost Arrange the decision tree such that there is only 3 branches on the most common paths and the more likely branch is always fall-through. A profile guided optimization pass would probably do something similar.	2016-08-30 05:24:16 +03:00
Ari Koivula	82cfab58f8	Improve fast mvd coding cost estimation A lot of time is being taken up by this function on ultrafast, and it doesn't do a very good job. This change aims to both simplify the logic and make the estimate better. The logic is simplified by using a look up for the step mvd bit cost step function instead of mimicking the binarization process. The estimation is made better by checking fractional cabac bit costs. The new function returns the same results as kvz_get_mvd_coding_cost_cabac, but is also faster than the old function.	2016-08-30 04:55:09 +03:00
Ari Koivula	d31be8eb27	Make mvd_coding_cost functions take const cabac	2016-08-30 04:46:46 +03:00
Ari Koivula	64d631c174	Fix 8bit to 10bit input conversion regression	2016-08-25 22:09:40 +03:00
Ari Koivula	27789125d8	Fix input bit depth conversion The input was being shifted to the wrong direction.	2016-08-25 22:05:25 +03:00
Ari Koivula	4ec039004b	Add monochrome encoding Write bitstream without chroma when encoding with --input-format=P400. This reduces bitstream size by 0-1 %, compared to coding monochrome in 420 format, and speeds up encoding slightly due to not processing chroma.	2016-08-25 20:15:26 +03:00
Ari Koivula	c5b70cf812	Add chroma format support to yuv_t	2016-08-24 19:20:53 +03:00
Ari Koivula	032ed30ff4	Add chroma format support to kvz_picture Add picture_alloc_csp to libkvz api to allocated pictures with chroma format different from 420.	2016-08-24 19:20:53 +03:00
Ari Koivula	48ccc26839	Add --input-format and --input-bitdepth Adds reading of 10 bit input for 10-bit encoding.	2016-08-24 19:20:53 +03:00
Ari Koivula	cc08073615	Refactor some indexing weirdness in init_lcu_t I thought there might be a bug in this so I cleaned it up.	2016-08-24 19:12:48 +03:00
Ari Koivula	b6d674d66e	Refactor integer vector inter prediction This code was pretty bad, so I cleaned it up a bit.	2016-08-24 19:09:26 +03:00
Ari Lemmetti	28c4174d0e	Fix incorrect shuffle parameters _MM_SHUFFLE uses reverse order	2016-08-23 19:40:46 +03:00
Ari Lemmetti	ce77bfa15b	Replace KVZ_PERMUTE with _MM_SHUFFLE The same exact macro already exists	2016-08-22 19:08:46 +03:00
Jovasa	68eef660bd	Fixed search around mv_in in fullsearch not being saved.	2016-08-19 15:19:29 +03:00
Eemeli Kallio	99d8b9abeb	Changed skip_rdoq name to kvz_skip_unnecessary_rdoq. Changed the order it uses when it goes through CGs and tuned its sum calculation.	2016-08-18 14:02:56 +03:00
Eemeli Kallio	1fb4755f31	Added rdoq-skip to quant-generic.c	2016-08-18 12:17:54 +03:00
Eemeli Kallio	d20ac03ca2	Added --rdoq-skip option	2016-08-18 12:17:53 +03:00
Marko Viitanen	83cf801664	Fixed MV constraint condition in bipred	2016-08-18 08:53:17 +03:00
Marko Viitanen	5ae1c595f2	Fixed slice_temporal_mvp_enabled_flag and disabled TMVP with tiles - slice_temporal_mvp_enabled_flag should be signalled also with non-IDR I-slices	2016-08-10 14:51:41 +03:00
Marko Viitanen	5326519182	TMVP cleanup and const qualifier fixes	2016-08-10 14:10:43 +03:00
Marko Viitanen	f40907260d	Added config parameter for TMVP and cmdline option --no-tmvp - Enabled by default - Cannot be used with GOP at the moment	2016-08-10 14:09:29 +03:00
Marko Viitanen	fd52dac1f7	Fixed TMVP scaling	2016-08-10 14:09:28 +03:00
Marko Viitanen	c664bc8cf7	Added flag collocated_ref_idx to the slice header	2016-08-10 14:09:28 +03:00
Marko Viitanen	c5f2611a38	Fixes for TMVP to work with the new CU array	2016-08-10 14:09:28 +03:00
Marko Viitanen	d85af5755b	TMVP working when only 1 ref frame	2016-08-10 14:09:28 +03:00
Marko Viitanen	39f0165efe	Fix a bug in TMVP, the reference cu_array was being overwritten	2016-08-10 14:09:27 +03:00
Marko Viitanen	adab8c327e	Clean TMVP code	2016-08-10 14:09:20 +03:00
Marko Viitanen	5fa8226ac9	Temporal merge candidate selection	2016-08-10 14:09:20 +03:00
Marko Viitanen	f83042f4a1	Temporal MV candidate selection	2016-08-10 14:09:19 +03:00
Marko Viitanen	f8671581e3	Implemented function kvz_inter_get_temporal_merge_candidates()	2016-08-10 14:09:19 +03:00
Marko Viitanen	2956bdb379	Added flag slice_temporal_mvp_enabled_flag	2016-08-10 14:09:19 +03:00
Arttu Ylä-Outinen	2a946bd88e	Rename encoder_state_t.global to frame "Frame" is more accurate than "global" since when OWF is used, encoder states for each frame have their own struct.	2016-08-10 13:22:36 +09:00
Arttu Ylä-Outinen	5fbb0a8c27	Fix includes	2016-08-10 13:05:40 +09:00
Arttu Ylä-Outinen	aabf6ca3ee	Extract encoding code from encoderstate.c Moves functions kvz_encode_coding_tree and kvz_encode_coeff_nxn from encoderstate.c to encode_coding_tree.c.	2016-08-09 22:16:50 +09:00
Arttu Ylä-Outinen	803f29be8f	Remove reconstructed picture allocation in lossless. Changes encoder_set_source_picture to set the reconstructed picture to a copy of the source picture instead of allocating a new picture when lossless coding is used.	2016-08-03 14:25:08 +09:00
Arttu Ylä-Outinen	aaec473a19	Refactor encoder state initialization. - Moves allocation of the reconstructed picture after the source picture is set. - Extracts main state initialization to a separate function from encoder_state_new_frame. - Changes kvz_encoder_feed_frame to return the frame. - Renames some functions to better match their purpose.	2016-08-03 14:25:08 +09:00
Arttu Ylä-Outinen	cd7024b3a5	Skip computing SSD when using lossless coding. The SSD is always zero since it is lossless.	2016-08-03 14:25:08 +09:00
Arttu Ylä-Outinen	fbbe5d1844	Use kvz_pixels_calc_ssd for SSD in search.c. Replaces loops for computing SSDs by calling kvz_pixels_calc_ssd in search.c.	2016-08-03 14:25:08 +09:00
Arttu Ylä-Outinen	22cc97ffb1	Fix missing field initializers.	2016-08-03 14:25:08 +09:00
Arttu Ylä-Outinen	06b82bf888	Disable filters, trskip and signhide in lossless. When lossless coding is used, deblock and SAO are skipped, transform skip flag is not written and sign hiding is not used.	2016-08-03 14:25:08 +09:00
Arttu Ylä-Outinen	97451ec401	Align assignments in encoder.c.	2016-08-03 14:25:08 +09:00
Arttu Ylä-Outinen	1dc94663c3	Bypass transform and quantization with --lossless. When --lossless is given, set cu_transquant_bypass_flag for every CU and bypass transform and quantization by directly copying reference pixels to reconstruction and the residual to coefficients.	2016-08-03 14:25:08 +09:00
Arttu Ylä-Outinen	2113b0182d	Enable PPS-level tq bypass flag with --lossless. Sets transquant_bypass_enable_flag to true in PPS when --lossless is given.	2016-08-03 14:25:08 +09:00
Arttu Ylä-Outinen	a5897bbece	Make cabac context initialization tables static.	2016-08-03 14:25:08 +09:00
Arttu Ylä-Outinen	23e7d9bb37	Add --lossless command line parameter.	2016-08-03 14:25:08 +09:00
Arttu Ylä-Outinen	5372ea432f	Update README and manpage.	2016-08-03 14:25:08 +09:00
Ari Lemmetti	6bcba004ff	Comment out to fix unused code error on clang.	2016-07-14 14:12:16 +03:00
Ari Lemmetti	c0979ebdcb	Implement AVX2 luma sampling	2016-07-14 12:53:02 +03:00
Ari Lemmetti	6244560426	Add avx2 strategy for kvz_filter_frac_blocks_luma.	2016-07-14 12:53:02 +03:00
Ari Lemmetti	9c4e9e049b	Load only what is needed. Eliminate latency from hadds.	2016-07-14 12:53:01 +03:00
Ari Lemmetti	7f71cb423a	Check 4 fractional pixel positions simultaneously	2016-07-14 12:52:24 +03:00
Ari Lemmetti	ad445ab8a1	Transition to kvz_filter_frac_blocks_luma	2016-07-14 12:51:02 +03:00
Ari Lemmetti	fccfbd2f28	Add strategy for kvz_filter_frac_blocks_luma	2016-07-14 12:51:02 +03:00
Ari Lemmetti	e9c3074d32	Add buffers and definitions for upcoming filtering Samples are to be filtered in separate blocks instead of making one big picture with interpolated pixels	2016-07-14 12:51:02 +03:00
Ari Lemmetti	7afe7e963b	Use fme_level to control the search accuracy.	2016-07-14 12:51:01 +03:00
Ari Lemmetti	5fa323bf25	Skip searching best hpel twice. Make hpel and qpel loops similar.	2016-07-14 12:51:01 +03:00
Ari Lemmetti	bc98a9affa	Change the search order to suit lighter fme search	2016-07-14 12:51:01 +03:00
Ari Lemmetti	2b0c8db349	Add quad satd for avx2	2016-07-14 12:50:24 +03:00
Ari Lemmetti	0ff69fd6f8	Add any size multi satd	2016-07-14 12:48:37 +03:00
Ari Lemmetti	d17b9e7d6e	Allow subme parameters 0-4 Update usage, presets,defaults,lib version	2016-07-12 19:49:38 +03:00
Arttu Ylä-Outinen	62ad57d0bf	Fix kvz_image_list_add for zero-sized lists. When a list does not have space for the new element, its size is doubled. If the size of the list is zero, it would not be resized. Fixed to always resize the list so that the new element can be added.	2016-06-22 13:35:16 +09:00
Arttu Ylä-Outinen	433e528af7	Drop unused variable in search_pu_inter. Removes unused variable max_px_below_lcu.	2016-06-22 13:35:16 +09:00
Arttu Ylä-Outinen	7836ff6ec9	Drop unused functions. Removes functions kvz_coefficients_calc_abs, kvz_intra_rdo_cost_compare and kvz_rdo_cost_intra which are no longer used.	2016-06-22 13:35:15 +09:00
Arttu Ylä-Outinen	e4b5840f56	Add parentheses around macro arguments in cabac.h.	2016-06-22 13:35:15 +09:00
Arttu Ylä-Outinen	a387b74e51	Fix resolution auto-detection. Only try to guess the resolution from filename when neither width nor height is given.	2016-06-22 13:35:15 +09:00
Arttu Ylä-Outinen	097bf8f3c0	Add a typedef for mvd coding cost functions.	2016-06-20 13:56:10 +09:00
Arttu Ylä-Outinen	d3c0e49286	Update comments.	2016-06-16 20:25:08 +09:00
Arttu Ylä-Outinen	ae832cda8c	Pack cbf flags in cu_info_t to two bytes. Reduces size of cu_info_t.	2016-06-16 20:24:19 +09:00
Arttu Ylä-Outinen	cad2d496b8	Enable 4x8 and 4x16 partition modes Enables search for 2NxN and Nx2N partition modes for 8x8 CUs and 2NxnU, 2NxnD, nLx2N and nRx2N partition modes for 16x16 CUs. Changes the loop for copying reconstructed luma pixels in kvz_inter_recon_lcu to use 4 byte chunks instead of 8 byte chunks since it is now possible to have 4 pixel wide blocks.	2016-06-16 20:23:16 +09:00
Arttu Ylä-Outinen	90df7350f0	Make deblocking work with 4 pixel wide blocks.	2016-06-16 20:21:50 +09:00
Arttu Ylä-Outinen	bf26661782	Add support for 4x4 blocks to SATD_ANY_SIZE. Makes functions satd_any_size_generic and satd_any_size_8bit_avx2 work on blocks whose width and/or height are not multiples of 8.	2016-06-16 18:53:17 +09:00
Arttu Ylä-Outinen	2ae260e422	Change width of cells in lcu_t to 4 pixels. Intra mode info for NxN partition units is now stored in the corresponding 4x4 cell in lcu_t.cu array.	2016-06-16 18:53:17 +09:00
Arttu Ylä-Outinen	360f5bb8da	Always use pixel coordinates for indexing lcu_t. Removes macro LCU_GET_CU and uses LCU_GET_CU_AT_PX in its place.	2016-06-16 18:53:17 +09:00
Arttu Ylä-Outinen	46e8122d27	Add functions for indexing cu_array_t structures. Replaces macro CU_ARRAY_AT with functions kvz_cu_array_at and kvz_cu_array_at_const.	2016-06-16 18:52:19 +09:00
Arttu Ylä-Outinen	c5afabdd3b	Change width of cells in cu_array_t to 4 pixels.	2016-06-15 12:25:11 +09:00
Arttu Ylä-Outinen	57a3d9b4b9	Add a function for copying CU data from LCUs. Adds function kvz_cu_array_copy_from_lcu which CU info data from an lcu_t structure to a cu_array_t structure.	2016-06-15 12:25:11 +09:00
Arttu Ylä-Outinen	2c85a00a55	Change kvz_cu_array_alloc to use pixel dimensions. Changes function kvz_cu_array_alloc to take width and height parameters in pixels instead of SCUs.	2016-06-15 12:25:11 +09:00
Arttu Ylä-Outinen	b276a347c0	Add a macro for indexing cu_array_t. Adds macro CU_ARRAY_AT(cu_array, x, y) to cu.h.	2016-06-15 12:25:11 +09:00
Arttu Ylä-Outinen	8ac1f1986e	Move CU array copy to a separate function. Moves code for copying parts of cu_array_t to a new function kvz_cu_array_copy in cu module.	2016-06-15 12:25:11 +09:00
Arttu Ylä-Outinen	41e75daed7	Fix overlapping memcpy in kvz_search_cu_smp. The destination and source pointers might be equal. Fixed by replacing the memcpy call with a simple assignment.	2016-06-15 12:25:11 +09:00
Ari Lemmetti	29af8bcd21	Remove const to match function signature	2016-06-14 18:19:40 +03:00
Eemeli Kallio	5af6ab320c	Merge branch 'me_early_terminate' Conflicts: configure.ac src/cfg.c src/cli.c src/kvazaar.h src/search_inter.c	2016-06-14 15:03:35 +03:00
Eemeli Kallio	43c7778b82	Updated version number.	2016-06-14 10:53:04 +03:00
Arttu Ylä-Outinen	23fdeeaf10	Move mv_cand and mv_dir into a bitfield in cu_info_t. Reduces size of cu_info_t.	2016-06-14 12:21:57 +09:00
Arttu Ylä-Outinen	35aadf6776	Reduce size of type in cu_info_t to two bits. Reduces size of cu_info_t.	2016-06-14 12:21:57 +09:00
Arttu Ylä-Outinen	1cbe844f79	Move inter and intra into an union in cu_info_t. Reduces size of cu_info_t.	2016-06-14 12:21:57 +09:00
Arttu Ylä-Outinen	b6d793ef33	Drop field inter.mvd from cu_info_t Instead of storing the mv differences in cu_info_t, they are computed from the mv candidates and the motion vector. Reduces the size of cu_info_t.	2016-06-14 12:21:57 +09:00
Arttu Ylä-Outinen	98aa906f30	Drop field coded from cu_info_t It can be inferred from the position and size of the CU.	2016-06-14 12:21:57 +09:00
Arttu Ylä-Outinen	ebb10763f1	Drop field inter.mv_ref_coded from cu_info_t. Storing inter.mv_ref_coded in cu_info_t is unnecessary since it can be computed from refmap and inter.mv_ref.	2016-06-14 12:21:57 +09:00
Arttu Ylä-Outinen	4be5c8f349	Move flags into a bitfield in cu_info_t. Reduces the size of cu_info_t.	2016-06-14 12:21:57 +09:00
Arttu Ylä-Outinen	30e9ee988d	Move bitcost field out of cu_info_t.inter. The bitcost is only needed for the currently searched CU. Fixes bitcost of the second PU being ignored when using SMP or AMP.	2016-06-14 12:21:57 +09:00
Arttu Ylä-Outinen	16d13ed046	Move cost field out of cu_info_t.inter The cost is only needed for the currently searched CU.	2016-06-14 12:20:05 +09:00
Arttu Ylä-Outinen	c5c2c182d9	Drop unused field mode from cu_info_t.inter.	2016-06-14 12:18:17 +09:00
Eemeli Kallio	e4f1a74512	Added early termination option for motion estimation. Conflicts: src/search_inter.c	2016-06-13 16:20:35 +03:00
Wassim Hamidouche	5bc7287c67	add fix for crypro	2016-06-09 10:49:31 +03:00
Wassim Hamidouche	35634b5596	correct MV sign encryption	2016-06-09 10:49:31 +03:00
Wassim Hamidouche	15abdc6e81	correct sign encryption	2016-06-09 10:49:31 +03:00
Wassim Hamidouche	73c3203a26	encry coef transfs	2016-06-09 10:49:31 +03:00
Wassim Hamidouche	7ad5f8bbe5	encry coef transf sign	2016-06-09 10:49:31 +03:00
Wassim Hamidouche	02b0712973	fix g++ compilation	2016-06-09 10:48:44 +03:00
Ari Koivula	a2170f0763	Compile the cryptopp wrapper only when used This should allow us to avoid an unnecessary dependancy to a C++ compiler. Conflicts: configure.ac	2016-06-07 17:11:12 +03:00
Ari Koivula	182038c743	Don't allow enabling encryption when it's not compiled in	2016-06-07 16:58:09 +03:00
Ari Koivula	8eb087120e	Make VisualStudio ignore the crypto stuff Add stubs for the crypto functions so we can refer to them, even if we never use them.	2016-06-07 16:58:09 +03:00
Wassim Hamidouche	76cb6dc6c2	add check flags	2016-06-07 10:54:26 +02:00
Ari Koivula	60ea8a359f	Add --crypto parameter	2016-06-07 10:31:40 +02:00
Wassim Hamidouche	02308d1ba6	add MVs encryption	2016-06-07 10:28:30 +02:00
Wassim Hamidouche	4637c8a828	compile Kvazaar encoder with ITpp library	2016-06-07 08:33:04 +02:00
Eemeli Kallio	8f182ac6de	Added functions select_starting_point and mv_in_merge to search_inter.c	2016-06-06 17:16:04 +03:00
Ari Koivula	fe71638a96	Fix problem with ASM compilation When compiling C++ files along with C, libtool would complain about the --tag missing, even though CC should be the default.	2016-06-06 15:47:56 +03:00
Eemeli Kallio	836a3b1daa	Added functions select_starting_point and mv_in_merge.	2016-06-06 12:18:33 +03:00
Ari Koivula	4eaacbe23e	Fix bug with lp-gop and ratecontrol The first frame was always qp51 due to gop_offset being -1 for the first frame. This fix makes it so that bits are allocated as if it was the last (high quality) frame from the previous GOP.	2016-05-27 15:53:55 +03:00
Ari Koivula	3fbd7ed97f	Add GOP layer weights for lowdelay-P When using ratecontrol with lowdelay-P, this improves BDRate by 1-25%. Strongest effect is when using 4 layers and multiple references. Also allow using 1 or 2 layers with ratecontrol.	2016-05-27 13:46:26 +03:00
Ari Koivula	67acead4bc	Fix referring over IDR boundary when using --gop This problem resulted in an illegal bitstream with --gop=lp, because it uses IDR's. The --gop=8 would not code IDR pictures, even when told to with -p, which masked this problem. This fix solves the problem with --gop=lp and also prevents references across the intra picture in --gop=8. The intra pictures should be set to IDR in a later fix, or an alternate method of differentiating between IDR and non-IDR intra should be made.	2016-05-27 13:20:53 +03:00
Ari Koivula	a77dc1610e	Refactor encoder_state_remove_refs I needed to debug this, so I rewrote it to make sense. There is an obvious bug with the IDR handling that I left in place to fix in a separate commit.	2016-05-27 13:20:45 +03:00
Eemeli Kallio	b5c05e58e0	Fixed typo in strategyselector.c	2016-05-24 11:04:29 +03:00
Ari Lemmetti	68c6f0f7b8	Enable deblocking for every preset Deblocking adds very little complexity while giving massive coding performance boost	2016-05-17 18:50:31 +03:00
Ari Lemmetti	6a07761b46	Add smp and amp options to presets	2016-05-17 14:26:58 +03:00
Ari Lemmetti	3107a93eaf	Fix avx2 chroma sampling for amp	2016-05-17 14:09:57 +03:00
Ari Koivula	24d0f9f685	Fix usage message for --hash	2016-05-11 15:03:43 +03:00
Ari Koivula	a1c772b696	Merge pull request #136 from MrAsura/cu-split-termination Cu split termination Closes #133.	2016-05-10 17:22:08 +03:00
Jaakko Laitinen	7010526b1d	Removed tabs.	2016-05-10 15:52:44 +03:00
Jaakko Laitinen	a77eb5c874	Fixed type conversion error when parsing cu split termination.	2016-05-10 14:34:46 +03:00
Jaakko Laitinen	0d361d5bc7	Moved cu split termination from a pre-processor to a input parameter.	2016-05-10 14:15:41 +03:00
Ari Koivula	1dbe4eb852	Merge branch 'mv-full'	2016-05-10 13:28:07 +03:00
Ari Koivula	f6a9d237a3	Merge pull request #134 from miimiz/testink_eemeli Strategyselector prints	2016-05-10 13:27:23 +03:00
Eemeli Kallio	8cfeed852c	Added print about SIMD optimizations available and in use to strategyselector.	2016-05-10 12:59:15 +03:00
Ari Koivula	f51a68b6fa	Add different sizes of search window for full search	2016-04-21 15:11:35 +03:00
Ari Lemmetti	efbdc5dade	Utilize registers more efficiently for 8x8 and larger blocks	2016-04-21 13:26:38 +03:00
Ari Lemmetti	192cee95b2	Vectorize vertical filtering	2016-04-21 13:26:38 +03:00
Ari Lemmetti	0be35f72b8	Filter 4 pixels simultaneously in x direction	2016-04-21 13:26:38 +03:00
Ari Lemmetti	10484bda9f	Make strategies out of fractional pixel sample functions	2016-04-21 13:26:38 +03:00
Ari Koivula	28e7548387	Fix bug in full mv search This optimization led to some points not being searched.	2016-04-21 12:03:57 +03:00
Ari Koivula	2576aeee0b	Use merge candidates in full mv search Perform a full search window around every mv candidate and the 0-vector.	2016-04-20 20:47:11 +03:00
Ari Lemmetti	8247faf8e0	Remove 64-bit only instruction to fix 32-bit compilation.	2016-04-19 18:05:11 +03:00
Ari Lemmetti	eb55d6b6b9	Fix writing over boundary.	2016-04-19 16:03:43 +03:00
Ari Lemmetti	bcabc6fadd	Remove pixel blit from strategies. Use memcpy instead.	2016-04-06 18:44:04 +03:00
Ari Lemmetti	2140197ccc	Tidy up coeff blit function and use memcpy again. Give memcpy constants for fixed sizes to enable copying many bytes simultaneously.	2016-04-06 18:03:00 +03:00
Ari Koivula	08b4480d94	Re-add time.h include Include-what-you-use wants to include sys/time.h instead, or if I override it to include time.h it will remove the include completely.	2016-04-02 19:05:16 +03:00
Ari Koivula	61fc3e87ba	Run include-what-you-use fix_includes.py fix_includes.py The includes should make more sense now and not just happen to compile due to headers included from other headers. Used a modified version of IWYU. Modifications were to attribute int8_t and so on to stdint.h instead of sys/types.h and immintrin.h instead of more specific headers. include-what-you-use 0.7 (git:b70df35) based on clang version 3.9.0 (trunk 264728)	2016-04-01 17:46:55 +03:00
Ari Koivula	016810d982	Move COMPILE_ macro to global.h While these are only used for strategies, it's non-intuitive to have to include strategyselector.h in every file under strategies before including anything else.	2016-04-01 17:46:55 +03:00
Ari Koivula	8908d85d66	Change all relative includes to absolute	2016-04-01 17:46:44 +03:00
Ari Koivula	4876879b82	Add IWYU pragmas	2016-03-31 12:33:34 +03:00
Marko Viitanen	41a5f9bbbe	Fix filetime conversion to timespec	2016-03-24 10:08:11 +02:00
Ari Koivula	9139e169fe	Fix unnecessary waiting in main thread The main thread has to wait for the worker threads to finish. The pthread_cond_timedwait call used to accomplish this was given a relative instead of absolute time, which resulted in the call returning immediately, because the time had already passed. This removes the now unnecessary sleeps and fixes the time given to the pthread_cond_timedwait such that it now waits until a job finishes or 100ms have passed.	2016-03-23 22:23:04 +02:00
Ari Koivula	e23ed231fb	Fix race condition with owf and non-square motion partitions The OWF wpp limit code assumed square blocks, and as such did not work correctly when height != width. This changes the relevant code to consider both height and width.	2016-03-22 16:46:38 +02:00
Arttu Ylä-Outinen	d6a3e02f16	Fix calculating reference CU index in inter search Fixes a possible segfault when SMP or AMP blocks are used.	2016-03-22 12:55:58 +02:00
Ari Lemmetti	f4538ab474	Copy pixels more efficiently in lcu recon.	2016-03-18 20:10:03 +02:00
Ari Koivula	5b66578f71	Add kvz_ prefix to md5 functions The non kvz_ symbols were being exported in the static lib, which got caught by Travis tests.	2016-03-18 13:13:35 +02:00
Ari Koivula	4125218cfa	Add --hash=md5 Add md5 through extras/libmd5 taken from HM with BSD license. It's implemented as a generic strategy using the same interface as checksum, so we can write a SIMD version if it seems necessary.	2016-03-18 05:23:57 +02:00
Ari Koivula	883448b8fb	Add --hash parameter Allows decoded picture hash to be selected among none and checksum.	2016-03-18 05:20:15 +02:00
Ari Lemmetti	6d5f8e3aec	Define KVZ_COMPILE_ASM for the correct files. Enables asm strategies again.	2016-03-17 16:21:31 +02:00
Ari Lemmetti	e502292ba8	Remove old function	2016-03-16 20:18:55 +02:00
Ari Lemmetti	c6cc96f5ec	Optimize sao band ddistortion	2016-03-16 20:16:00 +02:00
Ari Lemmetti	ab577f476f	Optimize sao reconstruct color	2016-03-16 20:15:32 +02:00
Ari Lemmetti	48bfddf4ec	Optimize calc sao edge dir	2016-03-16 20:14:50 +02:00
Ari Lemmetti	ba69992941	Optimize sao edge ddistortion	2016-03-16 20:14:19 +02:00
Ari Lemmetti	941b6b3e27	Optimize calc eo cat	2016-03-16 20:13:30 +02:00
Ari Lemmetti	04fbb48a09	Add strategy for avx2. Copy generic functions there.	2016-03-16 20:13:15 +02:00
Ari Lemmetti	4e30a215d8	Create generic strategy for sao.	2016-03-16 20:11:15 +02:00
Ari Koivula	6f431e510c	Comment and tidy threadqueue_worker Carefully avoided making any changes to the logic.	2016-03-14 20:08:04 +02:00
Ari Koivula	1165ae2e1f	Increase --mv-constraint=frametimemargin margin Increase the margin to be 4 luma pixels to every direction.	2016-03-14 16:02:54 +02:00
Arttu Ylä-Outinen	0eda28ced6	Fix Visual Studio warnings Initialization of a struct with addresses of local variables generated warning C4221 in encmain.	2016-03-14 14:12:21 +02:00
Ari Koivula	e91ca74733	Refactor kvz_encode_last_significant_xy	2016-03-10 18:47:16 +02:00
Ari Koivula	1fc0e8076c	Format kvz_encode_last_significant_xy whitespace	2016-03-10 18:17:45 +02:00
Ari Koivula	df9a958ef2	Merge branch 'log2'	2016-03-10 18:16:41 +02:00
Ari Koivula	4112a4364d	Remove g_to_bits table	2016-03-10 15:59:51 +02:00
Ari Koivula	9fcfba637f	Remove duplicated inline functions	2016-03-10 15:28:31 +02:00
Ari Koivula	e27ec2cc53	Add kvz_math.h for common inline math functions Calling it just math.h would have prevented including system math.h.	2016-03-10 15:26:18 +02:00
Ricardo Constantino	c515796a21	Only use version prefix in kvazaar binary Fixes regression since `54f08f2` causing libkvazaar version checks to not work (i.e. pkg-config)	2016-03-09 16:13:59 +00:00
Arttu Ylä-Outinen	54f08f2bdb	Use output of git describe as version.	2016-03-09 15:04:29 +02:00
Ari Koivula	f8edf28161	Fix const qualifier warning Also set the warning to an error in VS.	2016-03-09 14:16:15 +02:00
Ari Koivula	b0c3ece31e	Fix race condition when deblocking is on but SAO is off Already suspected this yesterday, but didn't want to add the code to handle it before confirming that it's actually a problem. It is.	2016-03-09 14:02:46 +02:00
Ari Koivula	1671725c72	Fix non-determinism issue with OWF WPP margin The previous reasoning used deblocking and fractional motion estimation together to arrive at a margin of 4 pixels. This was wrong, and with either of these off, half pixel chroma interpolation could use pixels outside the intended region. Deblocking does not currently affect the margin needed.	2016-03-08 20:18:38 +02:00
Ari Koivula	674bfa14ce	Comment WPP deblocking and SAO I was a bit unclear about exactly what happens and when regarding SAO and deblocking when we do frame-parallel WPP parallelism, so I checked and commented the bits that were unclear to me.	2016-03-08 19:39:04 +02:00
Ari Koivula	aec152c953	Fix OWF mv restriction limit The check was done in regard to the wrong dimension, allowing the access to unfinished parts of the frame when coding multiple frames at the same time.	2016-03-08 17:12:43 +02:00
Ari Koivula	fda103aa7c	Refactor cfg->tiles_width_count and cfg->tiles_height_count Change code everywere so these actually mean "width count" and not "width count minus one".	2016-03-07 17:29:15 +02:00
Ari Koivula	a350eb3a1e	Fix --tiles to have the correct number of tiles. The tiles_width_count etc. actually mean "count minus one".	2016-03-07 17:24:31 +02:00
Ari Koivula	49ea2d7b7f	Fix --mv-constraint=frametile Option --mv-constraint=frametilemargin was being used instead of frametile.	2016-03-07 16:41:00 +02:00
Ari Koivula	95b8dd99f6	Add --tiles parameter Add new parameter --tiles that accept only uniform split. I considered supporting the syntax of --tiles-width-split for this, but writing --tiles=u2xu2 is just not as intuitive as --tiles=2x2, and there is hardly ever any reason to use anything but uniform split. The more cumbersome --tiles-width-split and --tiles-height-split parameters are still there to allow finer control.	2016-03-07 16:33:51 +02:00
Ari Koivula	fd34dd9bc6	Fix race condition with OWF There was an off by one error in the dependance setting code, which resulted in dependencies not being set resulting in checksum errors. For example if ref_neg=1 and owf=1.	2016-03-07 13:38:23 +02:00
Ari Koivula	81b439f4da	Optimize starting point selection in tz Avoid checking zero motion vectors multiple times. The merge candidate list often has only one or two candidates, the other being zeroes.	2016-03-04 16:48:46 +02:00
Ari Koivula	2436702c27	Optimize starting point selection in hexbs Avoid checking zero motion vectors multiple times. The merge candidate list often has only one or two candidates, the other being zeroes.	2016-03-04 16:48:12 +02:00
Ari Koivula	5327b59b45	Remove KVZ_PERF_SEARCHPX It's too invasive and we don't really need it.	2016-03-04 16:48:12 +02:00
Arttu Ylä-Outinen	348ac4888b	Fix calc_mode_bits. The CUs left and above the current one would be set to NULL when there was only one CU between the current one and the left or top edge of the frame.	2016-03-04 14:08:35 +02:00
Ari Koivula	86219aa0fc	Fix non-determinism with tiles Earlier fix that fixed the supply side of the cu_array to take tile coordinates into account should have been accompanied with this one that does the same thing to demand side.	2016-03-03 17:39:20 +02:00
Arttu Ylä-Outinen	626b53ce85	Move sao search from encoderstate to sao. Moves sao search from function encoder_state_worker_encode_lcu in encoderstate.c to function kvz_sao_search_lcu in sao.c. Makes functions kvz_init_sao_info, kvz_sao_search_chroma and kvz_sao_search_luma static since they are no longer used outside sao.c.	2016-03-01 14:56:16 +02:00
Ari Koivula	cfa722e448	Reduce parallelism for tiles There is still some race-condition with encoding tiles from multiple frames, so disable this to keep the bitstream deterministic.	2016-02-29 20:20:21 +02:00
Ari Koivula	3dcc0957f8	Deal with impossible mv constraints If 0,0 vector is illegal, it's possible that no legal movement vector, is found, in which case a large cost is returned instead. The cost overflowed and there is all sorts of silliness with converting from double to int, but I'm not going to fix all of it because when we remove the doubles it will all get fixed.	2016-02-29 19:18:14 +02:00
Ari Koivula	b1adf1576a	Add --mv-constraint=frametilemargin Add an even stricter motion vector constraint to prevent motion vectors to fractional pixel positions that would need pixels outside the tile.	2016-02-29 19:18:14 +02:00
Ari Koivula	f808cbf608	Allow increased parallelism for tiles When movement vectors are constrained to tiles, only the same tile in previous frame needs to be depended upon.	2016-02-29 14:33:06 +02:00
Ari Koivula	f4ebff12b0	Combine tile mv constraint with OWF mv constraint This also fixes movement vectors in tiles when OWF is on. The OWF mv constraint assumed WPP, so it didn't work with tiles.	2016-02-29 14:33:06 +02:00
Ari Koivula	7981609cd0	Add --mv-constraint=frametile	2016-02-29 14:33:06 +02:00
Ari Koivula	9dbbb7fdbc	Add --mv-constraint argument	2016-02-29 14:33:06 +02:00
Ari Koivula	1be877faf9	Fix chroma reconstruction with tiles An incorrect frame boundary check caused a checksum error, because the chroma reconstruction of the encoder was wrong. The encoder treated horizontal tile boundaries as frame boundaries when the vertical component of the movement vector was a multiple of 8.	2016-02-29 14:32:51 +02:00
Ari Koivula	c0dc490dd1	Fix inter non-determinism with tiles CU data was being copied to the wrong place in the reference frames cu_array, which led to uninitialized data being used as a starting point for motion vector search. Fixes #99.	2016-02-26 17:05:04 +02:00
Ari Koivula	719d72925b	Add loop-input option This option is useful for testing long encodes, as you don't have to find an actual infinite input.	2016-02-18 20:00:55 +02:00
Ari Koivula	d23a5a15f1	Fix overflow in rate control A 32 bit int overflowed after 2^31 bits (2Gb). It will still overflow eventually, after 500 years of outputting 1Gb/s, but by that time, I recon we will have fixed this properly and it's time to upgrade.	2016-02-18 16:48:21 +02:00
Ari Koivula	eeafe14946	Clean up search initialization Copy lcu explicitly instead of initializing with the same parameters.	2016-02-17 14:57:31 +02:00
Arttu Ylä-Outinen	e5c84c361c	Eliminate a race condition with input thread. Changes communication between the input thread and main thread in encmain.c so that only one of them uses img_in and retval at a time. Fixes a race condition which would sometimes result in a deadlock.	2016-02-17 12:09:19 +02:00
Ari Koivula	c40ede56ad	Allow more frame parallelism in LP-gop Add dependency to the reference frame instead of the previous frame, in order to allow more frames to be encoded in parallel when temporal stepping >1 in LP-gop (such as --gop=lp-g8d4r1t2).	2016-02-05 17:08:24 +02:00
Arttu Ylä-Outinen	40c7198f7d	Add a script for updating README Adds script tools/update_readme.sh for regenerating the "Using Kvazaar" section of README.md from the output of "kvazaar --help".	2016-02-05 16:21:39 +02:00
Arttu Ylä-Outinen	aac5373095	Fix typos in documentation Fixes a few typos in README and command line help.	2016-02-05 16:21:27 +02:00
Ari Koivula	a4915dc547	Update man and README	2016-02-04 14:16:58 +02:00
Ari Koivula	e941e21cd6	Enable errors about non-existing CLI options Set opterr and optind to their normal default values.	2016-02-04 13:48:58 +02:00
Ari Koivula	7a4bf94a52	Add --version and --help Also don't print help by default, because it's too long. Print a shorter usage message instead.	2016-02-04 13:48:48 +02:00
Ari Lemmetti	99e37ec235	Update old pixel type to the current one	2016-01-30 19:33:09 +02:00
Ari Koivula	c76a0951cf	Change version to 0.8.3	2016-01-28 21:21:02 +02:00
Ari Koivula	cb2121b1aa	Double time scale when field coding is used	2016-01-28 21:04:52 +02:00
Ari Koivula	8ad7d2a714	Move interlacing stuff to libkvazaaar API This moves the interlacing from CLI code to api->encoder_encode, in order to make it possible to use field coding through the lib API. The field order is now determined per frame, as FFmpeg gives it per frame and it's signaled per frame. As a side effect, the CLI also now prints info from frames instead of fields. While we might want to extend the API in the future to allow printing of more detailed information about fields, for now it's more important that the CLI uses the real lib API. PSNR calculation for interlaced frames disabled until we have a way to avoid deinterlacing the frame when it's not necessary.	2016-01-27 15:29:45 +02:00

... 6 7 8 9 10 ...

2476 commits